{"id":"d8d49a41-182f-4c9f-84a7-946461830b07","arxiv_id":"2505.11141","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new bilingual benchmark with per-question human accuracy and common mistakes shows current multimodal AI models still underperform humans on reasoning.","lead":"This paper releases a new benchmark of almost 10,000 bilingual reasoning questions, each tagged with how often humans answer correctly and which wrong answer they pick. It tests 11 multimodal AI models against human performance and finds they lag far behind, especially on visual reasoning.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Human accuracy metadata provenance is unverified and the 'average human' comparison is unweighted; the alignment conclusions rest on this unvalidated pillar.","rationale":"The reader's conditional verdict and weakest assumption point to the same load-bearing concern: the human accuracy rates are the benchmark's defining feature, and every alignment analysis—difficulty buckets, error-consistency, and the 'approached average human' headline—depends on them. The paper gives no evidence about how many respondents produced each rate, whether the online platforms are reliable, or whether the aggregated rates are stable. Without that, the benchmark's most distinctive contribution is unverified. I also note a secondary issue: the 'fake reasoning' conclusion in Section 4.4 lacks a neutral-text control and is not fully supported by Table 4, since human-solution additions improve GPT-4o and QvQ. This does not undermine the resource's potential, but it does reinforce that the strongest interpretive claims should be presented as hypotheses pending validation. The proposed independent human re-scoring of a stratified sample would directly settle the provenance question; if it confirms the scraped rates, the benchmark's central claims become far more credible, and the conditional verdict could move toward acceptance.","tokens_in":29383,"tokens_out":5502,"duration_ms":55517,"concrete_test":"Select a stratified random sample of ~200 questions (50 per category, balanced by human difficulty). Recruit independent human participants (e.g., 50 per question on a controlled crowdsourcing platform) and compute per-question accuracy and modal wrong option. Compare with the scraped metadata using rank correlation and mean absolute error; also compare the unweighted 68.82% average with a response-count-weighted average using the platforms' underlying counts. If correlation is low or if the weighted human average differs by more than ~2 points, the difficulty stratification and model-vs-human alignment conclusions (including 'approached average human') should be revised or re-validated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 (Data Collection) states that human success rates and error-prone options were scraped from civil service exam prep platforms, but gives no per-question response counts, platform quality checks, or independent verification. These rates are used as ground-truth difficulty labels throughout Section 4.3 and to compute the 'average human performance' of 68.82% in Appendix A.1. If scraped rates are biased (e.g., self-selected test-takers, platform-specific answer-key errors, variable sample sizes), then the difficulty buckets, error-consistency analyses, and the headline that Gemini-2.5-pro 'has already approached the average human performance' are not trustworthy. Moreover, the 68.82% figure is a simple mean of question-level rates; with heterogeneous response counts it is not a valid estimate of average human performance. The benchmark's central novelty—per-question human alignment metadata—is exactly the unvalidated component. The 'fake reasoning' claim (Section 4.4) is additionally under-specified: no control for adding a neutral passage of matched length, and Table 4 shows GPT-4o and QvQ improve under human solutions (+3.38, +0.43), contradicting the text's blanket 'degraded except Gemini' statement.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Human-Aligned Bench, a 9,794-question bilingual benchmark drawn from Chinese civil service examinations across four reasoning categories (visual reasoning, definition judgment, analogical reasoning, logical judgment). Each question is accompanied by metadata consisting of a human correctness rate, the human error-prone option, and category-level human solution frameworks. The authors evaluate eleven proprietary and open MLLMs and report three main findings: current MLLMs are far from human performance on visual reasoning; MLLMs do not show human-like accuracy trends across human difficulty bins in visual reasoning; and adding human or self-generated solution summaries to prompts degrades most models, which the paper interprets as evidence of 'fake reasoning'. An additional claim is that Gemini-2.5-pro-exp-03-25 has approached the average human performance on the benchmark.","tokens_in":29628,"tokens_out":6639,"duration_ms":64393,"significance":"If the human metadata is validated, the benchmark would fill a real gap: most multimodal reasoning benchmarks lack fine-grained human performance data, error-prone distractors, and solution strategies, and the bilingual coverage is broader than that of MM-IQ, VisuLogic, and VISUALPUZZLES. The dataset and code are released, which supports reproducibility, and the evaluation covers a wide range of closed and open models. The error-consistency analyses in Figure 4 are a potentially distinctive contribution. However, the paper's central claims currently rest on an unvalidated source of human performance data and on an internally contradicted 'fake reasoning' analysis, so the contribution is not yet fully established.","major_comments":[{"comment":"The per-question human correctness rates and error-prone options are scraped from online civil-service exam-preparation platforms, but the paper reports no response counts per question, no platform identifiers, no inclusion/exclusion criteria, no inter-annotator agreement, and no independent validation against a held-out human sample. These rates are the ground truth for the difficulty bins in Table 3, for the accuracy-trend claims in Section 4.3, and for the 68.82% human baseline in Appendix A.1, so the headline that Gemini-2.5-pro-exp-03-25 has 'approached the average human performance' inherits any bias in this metadata. Moreover, 68.82% is a simple mean of question-level rates, which is a valid estimate of average human performance only if every question has the same number of respondents; the paper should report per-question response counts or use a response-weighted estimate.","section":"Section 3 (Data Collection) and Appendix A.1"},{"comment":"The claim that 'the performance of MLLM is degraded except for Gemini-2.5-pro-exp-03-25' is directly contradicted by Table 4, where GPT-4o-WHS improves by +3.38 overall and QvQ-72B-Preview-WHS improves by +0.43. The experimental design also has no control condition with a matched-length neutral passage, so the observed degradation, where it occurs, cannot be attributed specifically to human reasoning content rather than to prompt length, format, or distraction. The 'fake reasoning' conclusion should be reworded and supported with neutral controls, per-category results, and a statistical comparison.","section":"Section 4.4 (Fake Reasoning Analysis) and Table 4"},{"comment":"All results are reported from single runs at temperature 0.6, with no repeated sampling, no confidence intervals, and no significance tests. Consequently, fine-grained comparisons such as QvQ-Max outperforming o4-mini by 0.37% on Chinese questions (Section 4.2, Compare on Language Type) and the small differences among models in Table 3 cannot be distinguished from sampling noise. The authors should report multiple runs and variance estimates for the headline rankings and for the consistency analyses in Figure 4.","section":"Section 4.1 (Experimental Setup)"},{"comment":"The visual-reasoning per-bin accuracies are reported without per-bin question counts, and the 0-20% and 20-40% columns contain small numbers of questions, so a difference of one or two percentage points can correspond to a handful of items. The claim in Section 4.3 that models exhibit uniform or inverse accuracy trends on visual reasoning across human difficulty levels is therefore not statistically supported as reported. The table should show per-bin N and human mean accuracy per bin, and the monotonicity claim should be tested statistically rather than by visual inspection.","section":"Table 3 and Section 4.3 (Visual Reasoning rows)"},{"comment":"The benchmark questions are drawn from public civil service exams that are widely available online, and the paper reports no contamination check or discussion of the possibility that the evaluated models have seen these exact questions during training. Because the claims of approaching human performance and of 'fake reasoning' assume that models are solving rather than recalling, the authors should at least report exact-match or n-gram overlap analyses where feasible and discuss the residual contamination risk for closed models.","section":"Section 4 (Overall Results)"}],"minor_comments":[{"comment":"The comparison row labels the benchmark 'Fake Reasoning' rather than 'Human-Aligned Bench'; this appears to be a typo from an earlier name and should be corrected.","section":"Table 1"},{"comment":"The table uses the abbreviation 'WHS' for both 'With human Solution' and 'With self Solution'; the second should be renamed to 'WSS' for clarity.","section":"Table 4"},{"comment":"The text appears to swap the subfigures: Section 4.2 attributes the language-type comparison to Figure 3(a), while the figure caption assigns modalities to (a) and languages to (b).","section":"Section 4.2 and Figure 3"},{"comment":"The expression 'approximately 75% (performance percentage / 60)' is unexplained, and no equations define the response-consistency or error-consistency metrics, so Figure 4 cannot be reproduced from the text.","section":"Section 4.3 (Consistency on Response)"},{"comment":"The paper uses 'human correctness rate', 'human accuracy rate', and 'human score rate' interchangeably; the authors should use one term consistently and define it as the percentage of respondents who chose the correct answer.","section":"Throughout"},{"comment":"The abstract says the benchmark includes 'bilingual (Chinese and English) multimodal questions', but the English versions are GPT-4o translations of Chinese exam questions; this should be stated explicitly in the main text so that language comparisons are not misread as comparisons of original source languages.","section":"Abstract and Table 2"}],"recommendation":"major_revision","confidential_remarks":"The dataset and code release are valuable, and the benchmark has the potential to be a useful community resource. However, the provenance and statistical adequacy of the human performance metadata are the foundation of the paper's main claims, and the current description is insufficient. The 'fake reasoning' claim is also stated more strongly than the data in Table 4 support, and the text should be corrected to match the table. I recommend major revision and would want to see the human metadata validation and the corrected fake-reasoning analysis before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The benchmark is real: 9,794 bilingual multimodal questions from Chinese civil-service exams, four reasoning categories, plus per-question human correct rates and human error-prone options. That metadata combination is new relative to MM-IQ, VisuLogic, and VISUALPUZZLES. The paper also gives per-category human solution strategies, which makes a range of alignment analyses possible. If the data release matches the description, it is a resource worth having.\n\nWhat the paper does well: the dataset construction is described enough to reimplement, the categories are sensible, and the comparison to prior benchmarks is fair. The citation pattern is clean, with no self-citation padding. The basic empirical finding—that MLLMs are far stronger at text reasoning than visual reasoning and show non-monotonic accuracy across human difficulty levels in visual tasks—is interesting and plausible.\n\nSoft spots, in order of importance. The human success rates are scraped from online exam-prep platforms with no per-question response counts, no checks on platform quality, and no independent verification. That is exactly the load-bearing novelty: difficulty bins, error-consistency analysis, and the 68.82% 'average human performance' all come from these rates. A simple mean of question-level rates is not a valid estimate of average human performance when response counts are heterogeneous, so the claim that Gemini-2.5-pro has 'approached the average human' is not trustworthy yet.\n\nSecond, the fake-reasoning analysis in Section 4.4 overreaches. The text says performance degrades when human or self solutions are added, except Gemini. Their own Table 4 shows GPT-4o overall +3.38 with human solutions and QvQ +0.43. There is also no control condition with a matched neutral passage, so attributing the drop to 'fake reasoning' is unsupported. This is fixable, but as written it is a real, self-contained flaw.\n\nThe evaluation also runs one temperature with no repeats or error bars; that is minor for a benchmark paper but worth fixing.\n\nBottom line: the benchmark deserves a serious referee and likely a conditional-accept path, not a desk reject. The authors need to validate the human metadata, report response counts, and either redo or carefully qualify the fake-reasoning claim. I would bring it to reading group, and I would cite it once the data is out and the provenance questions are answered.","headline":"Genuinely useful new benchmark, but the load-bearing human-performance metadata is unvalidated and the 'fake reasoning' claim overreaches against its own Table 4.","tokens_in":30142,"tokens_out":3039,"would_cite":true,"duration_ms":29998,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that attaching human accuracy rates, common wrong answers, and solution strategies to 9,794 bilingual reasoning questions reveals that current multimodal models do not reason like humans, and that much of their apparent…","keywords":["multimodal large language models","reasoning benchmark","human alignment","visual reasoning","fake reasoning","civil service examination","bilingual evaluation"],"falsifier":"Take a stratified random sample of about 600 questions, have a fresh panel of human participants answer them under exam-like conditions, and compare the resulting accuracy rates with the scraped online rates; if the rates diverge substantially, especially on hard questions, the benchmark's difficulty ladder and human-alignment conclusions would need recalibration. Separately, if supplying models with human- or self-generated solution strategies consistently improves accuracy across model families, the paper's fake-reasoning claim would be contradicted.","tokens_in":29228,"feed_emoji":"🧠","tokens_out":7878,"duration_ms":77394,"temperature":0.7,"pith_summary":"The paper introduces Human-Aligned Bench, a set of 9,794 bilingual reasoning questions taken from Chinese civil service examinations, where every question comes with the human success rate, the option humans most often pick when wrong, and a written summary of how human experts approach that question type. The authors use this resource to test whether multimodal large language models reason the way humans do, rather than merely whether they answer correctly. They find that current models trail human accuracy overall, that their accuracy fails to track human-defined difficulty on image-based reasoning, and that injecting either human or self-generated solution strategies usually does not help, and often hurts, performance. The central conclusion is that much of what looks like reasoning in these models may be fake reasoning: it does not follow from the model's own understanding of the problem type.","feed_headline":"9,794-question benchmark shows AI reasoning isn't human-like","feed_subtitle":"Each question's human accuracy and common errors make model-vs-human comparison fine-grained.","key_machinery":"The central object is the Human-Aligned Bench dataset: 9,794 bilingual questions in four categories, visual reasoning, definition judgment, analogical reasoning, and logical judgment, each carrying a human correctness rate, a human error-prone option, and a per-category summary of human solution strategies. The dataset is assembled from Chinese civil service examination papers, whose questions are designed to require contextual reasoning with no outside knowledge, and it is translated into English to give bilingual coverage. The annotation layer is what carries the argument: it lets the authors measure accuracy in human-defined difficulty buckets, compute the consistency of model responses and model errors with human responses and errors, and test whether feeding models human or self-generated solution frameworks changes their accuracy. The solution frameworks themselves are the probe for fake reasoning.","core_discovery":"The paper's central claim is that a reasoning benchmark can be made human-aligned by attaching three human measurements to every question: the rate at which human test-takers answer correctly, the distractor option that humans most often choose when they are wrong, and a written account of how human experts solve that question type. Using 9,794 questions from China's civil service examination, chosen because they test pure contextual reasoning rather than specialized knowledge, the authors compare eleven multimodal models with these human measurements. They report that text-based reasoning roughly tracks the human difficulty gradient, while visual reasoning does not: model accuracy stays flat or moves against the gradient as questions get harder for humans. On the alignment measures, models rarely choose the same wrong options that humans choose on easier questions, though their errors become more human-like on the hardest questions. The paper interprets the solution-injection experiments as evidence that most current models do not reason from their own understanding of a question type, describing the phenomenon as fake reasoning.","pith_inferences":["If the online exam-prep accuracy statistics are not representative of the general test-taking population, the human baselines could overstate or understate true difficulty; a controlled replication on a fresh sample would settle this.","The fake-reasoning diagnosis could be sharpened by varying how strongly the injected solution is phrased, or by fine-tuning models on solution strategies to see whether the degradation disappears.","The same template could be applied to other standardized exams with published item-level statistics, turning any such exam into a human-aligned reasoning probe.","A stronger test of alignment would condition model errors on the solution strategy used; if models fail to make the same errors as humans even when given human strategies, the gap is in reasoning itself rather than in answer selection."],"forward_implications":["Model accuracy can be reported as a function of human-defined difficulty, so a flat or inverted difficulty curve becomes an explicit diagnostic for visual reasoning rather than a hidden artifact.","The benchmark gives a decomposition of model skill by language and modality, isolating whether a failure is perceptual, textual, or reasoning-based.","The fake-reasoning result implies that prompt sensitivity is a measurable property of reasoning models and should be reported alongside average accuracy.","The human error-prone option makes it possible to track whether models are led astray by the same distractors that mislead people, turning error analysis into an alignment metric."],"supporting_citations":[{"why":"Provides the closest prior benchmark for reasoning without domain knowledge that Human-Aligned Bench extends by adding human performance annotations.","marker":"[27]"},{"why":"Earlier multimodal reasoning benchmark used as a comparison baseline; lacks human accuracy and error-prone option metadata.","marker":"[28]"},{"why":"Visual reasoning benchmark used to position Human-Aligned Bench's wider coverage and fine-grained human alignment data.","marker":"[29]"},{"why":"A strong proprietary multimodal model whose results contribute to the overall performance comparison and the text-versus-image gap.","marker":"[36]"},{"why":"Model family whose best member approaches average human accuracy on the full benchmark, grounding the main human-alignment result.","marker":"[37]"},{"why":"Text-only reasoning model used as an LLM baseline to show that multimodal models still lag on pure text reasoning.","marker":"[17]"},{"why":"Open-source multimodal model family whose comparative results support the open-versus-proprietary performance discussion.","marker":"[46]"}],"fun_headline_variants":["AI's visual reasoning diverges from human difficulty curve","Benchmark: AI fakes reasoning on visual questions","9,794 questions expose AI's fake visual reasoning","Human-aligned test: AI's visual logic breaks down","AI reasoning benchmark: visual tasks don't match human difficulty"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the human accuracy rates scraped from online civil service exam preparation platforms accurately represent how real test-takers perform; if those rates are biased, then every difficulty label, every human-model alignment comparison, and the claim that the best model has approached average human performance would lose their foundation.","fun_headline_variants_meta":{"raw":{"variants":["AI's visual reasoning diverges from human difficulty curve","Benchmark: AI fakes reasoning on visual questions","9,794 questions expose AI's fake visual reasoning","Human-aligned test: AI's visual logic breaks down","AI reasoning benchmark: visual tasks don't match human difficulty"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000711,"raw_usage":{"total_tokens":3198,"prompt_tokens":938,"completion_tokens":2260,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":2182}},"tokens_in":554,"tokens_out":2260,"duration_ms":16320,"temperature":1.0,"reasoning_tokens":2182,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:56:22.724117+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a stratified random sample of about 600 questions, have a fresh panel of human participants answer them under exam-like conditions, and compare the resulting accuracy rates with the scraped online rates; if the rates diverge substantially, especially on hard questions, the benchmark's difficulty ladder and human-alignment conclusions would need recalibration. Separately, if supplying models with human- or self-generated solution strategies consistently improves accuracy across model families, the paper's fake-reasoning claim would be contradicted.","supporting_citations":[],"review_version":1}