{"id":"aee1d3a1-5d81-4877-adac-d396354208d4","arxiv_id":"2412.10056","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"On a fresh, non-leaking Gaokao benchmark, LLM accuracy stays nearly flat across question difficulty, diverging from human Rasch curves and suggesting high scores can overstate human-like ability.","lead":"This paper builds a new benchmark from China's 2024 college entrance exam and tests AI models released before the exam. It finds that model scores do not follow the human-like difficulty curve, so high scores may not signal genuine understanding.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Concern: The Rasch-model mismatch (Fig. 6, R²=-0.23) is asserted against an assumed human curve; with no human response data on the 2024 items and an aggregate fit that is not a valid Rasch fit, it cannot bear the central claim.","rationale":"The reader identified the missing human baseline as the weakest assumption; I agree, and I think it is the load-bearing issue. The paper's own stated evidence for the central claim is the Rasch comparison (Sec. 3.1), not the qualitative error examples. The negative R² has an additional internal problem: even if every LLM were perfectly Rasch-consistent, pooling across models and subjects with different abilities can produce aggregate scoring rates that are not logistic, so the poor aggregate fit is not a test of Rasch consistency. The paper does report a high correlation (0.94) between Elo difficulty and expert annotations, which supports the difficulty ordering as human-perceived; but human-perceived difficulty is not human response probability, and the Rasch model is specifically about response probabilities. Without item-level human response data, the central claim reduces to 'LLMs do not match a curve that we assume humans follow.' I also note that the models' raw scores (roughly 60–70% of total) are not obviously 'high', which further weakens the framing, but the missing human baseline is the decisive issue. The released benchmark and evaluation pipeline are real contributions, and the paper is transparent about what was done; the problem is that the analysis as written does not support the headline conclusion. Hence I keep the reader's REJECT rather than downgrading.","tokens_in":26025,"tokens_out":4596,"duration_ms":53068,"concrete_test":"Use the official Gaokao pretesting statistics or administer the same 2024 items to a representative cohort of high-school graduates under closed-book conditions; fit the Rasch model to the human item responses, estimate b_human, and plot human empirical item-characteristic curves. If human responses also yield R²≈−0.23 or high variance within difficulty bins, the Fig. 6 mismatch is not evidence about LLM capabilities. If humans fit Rasch well while LLMs do not, the concern is resolved. As a complementary analytical check, re-fit an item-response model with model-specific ability θ_j and item difficulty b_i directly to the released per-item scores; if the per-model Rasch fit is adequate, the aggregate negative R² is an artifact of pooling.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is 'high scores do not necessarily reflect human-aligned capabilities' (Sec. 3.1). Its main quantitative support is the poor fit of the Rasch model to LLM scoring rates (Fig. 6, R²=-0.23). For this to be evidence, two premises must hold: (1) human examinees on these 2024 Gaokao items follow the one-parameter Rasch curve with item difficulties b, and (2) the Elo-derived difficulties plotted on the x-axis are those human difficulties. Neither is established. No human response data on the 2024 items are reported; the 'human theoretical curve' is imported from cognitive psychology, and the Elo difficulties are a hybrid of expert annotation and LLM pairwise judgments, correlated with expert ratings (Fig. 5) but never checked against human response data. More fundamentally, the fitted curve in Fig. 6 is not a Rasch model fit. Rasch predicts P(X=1|θ,b)=1/(1+exp(b−θ)) for a fixed examinee ability θ, but the plotted points pool multiple models, subjects, and question types with different abilities; the aggregate of Rasch-consistent examinees need not be logistic, so a negative R² against this aggregate is not evidence of a human-alignment mismatch. The qualitative examples and grading-inconsistency findings are suggestive but do not substitute for the missing human baseline. The benchmark itself has strengths (temporal isolation, 54-teacher grading, released code), but the stated conclusion rests on the Rasch comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces GAOKAO-Eval, a 2024 Chinese Gaokao-based benchmark designed to be non-leaky and comprehensive, evaluates several LLMs released before the exam date, and reports that their scores do not reflect human-aligned capabilities. The main quantitative evidence is a Rasch-model comparison (Eq. 1, Fig. 6) showing a poor fit with R²=-0.23, alongside Pearson correlations between difficulty and scoring rates, variance computations within difficulty bins, teacher-grading inconsistency rates (Eq. 4), and qualitative examples of model errors. The authors also propose reasoning tokens as an LLM-aligned difficulty proxy, reporting an improved fit for o1 models in Fig. 11. The benchmark resource itself has strengths—temporal isolation, 54-teacher grading, and released code—but, as detailed below, the paper's central claim rests on a Rasch comparison made without human response data on the 2024 items and using an aggregate fitting procedure that is not a valid Rasch analysis.","tokens_in":26274,"tokens_out":5655,"duration_ms":63482,"significance":"If the Rasch mismatch were established against a genuine human baseline, the paper would make an important contribution to benchmark-validity research, showing that high benchmark scores need not imply human-like reasoning. The authors are explicit about their falsifiable prediction: LLM scoring rates should deviate from a human item-response curve. The benchmark construction and the release of code, model responses, and teacher-graded scores are concrete strengths that the community could build on. However, the stress-test concern raised in the review process is borne out by the manuscript: the human curve is assumed, not measured, and the pooled fit in Fig. 6 is not a Rasch fit in any standard sense. The paper's own limitation statement in Section 3.4 ('scores in GAOKAO-Eval should be interpreted with caution') further weakens the strength of the claims. For these reasons, the central conclusion is currently unsupported despite the value of the benchmark artifact.","major_comments":[{"comment":"The central Rasch comparison has no human response baseline. The text in Section 3.1 says 'we directly use this equation as the basis for evaluation' and the caption of Fig. 6 reports R²=-0.23, but no human examinee responses to the 2024 Gaokao items are collected or cited. The 'theoretical human performance curve' is imported from cognitive psychology (Rasch, 1993), while the x-axis difficulties come from the Elo system described in Section 3.1, which is only correlated with expert annotations (Fig. 5). Without human response data on the same items, a negative R² for LLMs does not establish a deviation from human-aligned capability; it only shows that the pooled LLM scoring-rate curve is not logistic. A concrete fix would be to estimate item difficulties and a human scoring-rate curve from pretesting data (which the Gaokao process nominally collects) and recompute the comparison.","section":"Section 3.1, Figure 6"},{"comment":"The fitted curve is not a valid Rasch fit. Eq. (1) defines the probability of a correct response for a single examinee with fixed ability θ, but the plotted points in Fig. 6 pool multiple models, subjects, and question types with different abilities. The mixture of Rasch-consistent examinees is not necessarily logistic, so fitting one logistic curve to the aggregate and reporting R²=-0.23 is not evidence against Rasch or IRT. The authors should instead fit a hierarchical or multilevel IRT model (e.g., random effects for models and subjects, or a 2PL model with a discrimination parameter) and report item-level fit statistics; alternatively, they should explicitly justify why the aggregate should preserve the logistic form.","section":"Section 3.1, Eq. (1) and Figure 6"},{"comment":"The Inconsistent Score Rate definition produces values near 32% for approximately normal score distributions, because P(|X−μ|>σ)≈0.317 for any normal distribution. Reporting 'over 32%' as 'too high' is therefore partly definitional without a human-grading baseline or a null model. The paper gives no ISR for human essays or human short answers graded under the same rubric, so the claim that LLM responses cause unusually high grading inconsistency is not established. This matters because the ISR discussion is used to support the 'high variance' finding in Section 3.2. A comparison against human responses under identical scoring conditions, or a simulation with calibrated teacher noise, is needed.","section":"Section 3.4, Eq. (4)"},{"comment":"The difficulty axis is not independent of the models being analyzed. Section 3.1 states that the hybrid difficulty ratings combine 'manual annotations with an Elo rating system' and that the system 'adjusts LLM scores based on pairwise comparisons,' and Fig. 5 shows Elo ratings derived partly from GPT-4o and GPT-4o-mini judgments. Because the same or similar LLM outputs are used to estimate item difficulties and to compute scoring rates, the Rasch comparison is partially circular. The reported 0.94 correlation with human expert annotations is a sanity check but does not replace human response data. A concrete test would be to re-estimate difficulties from human responses alone and recompute the Rasch fit with those difficulties.","section":"Section 3.1, difficulty estimation and Figure 5"},{"comment":"The paper's title and abstract claim that 'high scores' fail to reflect capability, but the evaluated models do not achieve high scores in an absolute sense. The top science total in Table 4 is 468.5/750, which is 62.5%, and most models are far below that. What the data actually show is that moderate scores are accompanied by a flat difficulty curve, not that near-ceiling scores fail to align with human difficulty. To support the stated claim, the paper would need to include a genuinely high-scoring model (with appropriate data-leakage controls) or substantially rephrase the title, abstract, and Section 3.1 to say that moderate scores do not imply human-aligned difficulty sensitivity.","section":"Section 3.1, Tables 4 and 5"}],"minor_comments":[{"comment":"The title contains a typo ('GAOKAO-E VAL') and the abstract has several grammatical and style issues, including 'phenomenons' and inconsistent capitalization of 'We' and 'we'; a careful proofreading pass is needed.","section":"Title and Abstract"},{"comment":"The text cites 'Query of CC technique (Fei et al., 2024)' but no Fei et al. entry appears in the reference list; either add the reference or remove the citation.","section":"Appendix A.1"},{"comment":"The 'semi difficulty-invariant' conclusion is based on Pearson correlations in Eq. (2), but no confidence intervals, significance levels, or per-cell sample sizes are reported, making it difficult to judge whether the near-zero correlations are stable estimates or noise.","section":"Section 3.2, Figure 8"},{"comment":"The o1 experiment is presented as evidence that reasoning tokens 'mitigate the mismatch,' but it involves a single model family, the R² values are still low (0.1019), and o1 models were released after the June 6, 2024 cutoff used elsewhere in the paper; the claim should be scoped accordingly.","section":"Section 4, Figure 11"},{"comment":"Section 2.2 says 54 teachers graded responses, while Section 3.4 says each question was reviewed by at least three teachers and the average was taken; the relationship between these statements is unclear and should be reconciled, including how the averaging affects the ISR calculation in Eq. (4).","section":"Sections 2.2 and 3.4"}],"recommendation":"reject","confidential_remarks":"The benchmark artifact has real value, and the authors are to be commended for releasing code and model outputs. However, the central scientific claim is not supported by the current analysis: the missing human response baseline and the invalid aggregate Rasch fit are load-bearing and cannot be repaired with local edits. A revised manuscript that reframes the contribution as benchmark construction, adds a proper IRT analysis with human response data, and uses a defensible high-scoring comparison would be worth reconsidering, but the present version is not publishable as a claim about high scores and human-aligned capabilities."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the benchmark itself is useful: 2024 Gaokao questions, evaluated only on models released before the exam, graded by 54 teachers, code and outputs released. That is a genuine resource. Second, the central claim—that high scores do not reflect human-aligned capabilities—is not established by the paper's own analysis. The Rasch comparison has no human baseline, and the ISR metric is essentially vacuous.\n\nWhat the paper does well: the temporal isolation is real, the dataset fills a gap, and the qualitative error examples (the fabricated poem, the correct answer via flawed steps) are compelling illustrations that benchmark scores can be misleading. The observation that o1's reasoning-token counts improve the IRT fit is a nice exploratory finding.\n\nThe soft spots are load-bearing. The models top out around 62%, which is not a 'high score' in the usual sense, so the title overstates the phenomenon. More importantly, the Rasch mismatch is measured against an assumed human curve, not against any human response data on these 2024 items. What is plotted is a pooled aggregate across models and question types; the aggregate of Rasch-consistent examinees need not be logistic, so a negative R² against that aggregate does not show a human-alignment mismatch. The Elo difficulties are themselves partly derived from LLM judgments, which adds circularity. And the ISR definition in Eq. 4—fraction of scores more than one standard deviation from the mean—is about 32% for any roughly normal distribution, so reporting 'over 32%' carries no information. The negative R² values lack confidence intervals, and the reasoning-token analysis is not statistically grounded.\n\nWhere does that leave things? The dataset is worth having and the paper is worth engaging with, but the headline conclusion needs reframing or a real human baseline. As is, the analysis does not support 'high scores do not reflect human-aligned capabilities' as a general claim.\n\nI would send this to a serious referee: the resource deserves review, and the flaws are fixable in revision, even if the central claim may need to be toned down.","headline":"The 2024 Gaokao dataset is the real contribution; the Rasch analysis anchoring the headline claim does not survive contact with the data.","tokens_in":26948,"tokens_out":2535,"would_cite":true,"duration_ms":26524,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"High Gaokao scores do not prove LLMs reason like humans","keywords":["LLM evaluation","Gaokao benchmark","Rasch model","question difficulty","human-aligned capability","data leakage","reasoning tokens","closed-book evaluation"],"falsifier":"Run the same 2024 Gaokao questions on a large sample of human examinees and fit the Rasch model to their scores; if the human data also show near-zero or negative R-squared, the reported mismatch is a property of the items or the difficulty scale, not evidence about LLM capabilities.","tokens_in":25729,"feed_emoji":"🎓","tokens_out":8037,"duration_ms":72277,"temperature":0.7,"pith_summary":"The paper builds a closed-book benchmark from China's Gaokao exams, uses human teachers to grade the responses, and evaluates only models released before the exam date. It argues that even when data leakage and benchmark coverage are controlled, high scores do not imply human-aligned capability. LLM scoring rates stay nearly flat as question difficulty rises, and they scatter widely for questions of similar difficulty. Using the Rasch model from psychometrics as a human-performance yardstick, the paper shows LLM scoring deviates strongly from the theoretical human curve, with a negative R-squared for the fit. A follow-on experiment with o1 finds that re-defining difficulty by reasoning tokens instead of human difficulty improves the fit, suggesting human-derived difficulty may be the wrong axis for LLMs.","feed_headline":"LLMs ace Gaokao yet ignore question difficulty","feed_subtitle":"A closed-book Gaokao benchmark finds LLM scores barely move as questions get harder, challenging high-score claims.","key_machinery":"The Rasch model, a one-parameter item response theory model, $P(X=1|\\theta,b)=e^{\\theta-b}/(1+e^{\\theta-b})$, supplies the human-performance reference curve. The paper pairs it with an Elo rating system, seeded by human expert annotations and LLM judgments, to assign difficulty values to Gaokao items; the Elo difficulties correlate with human expert ratings at up to 0.94. Two derived statistics carry the argument: the Pearson correlation between difficulty and scoring rate (near zero, giving 'semi difficulty-invariant' behavior) and the variance of scoring rates within small difficulty bins (high, violating the expected monotone decline). The o1 experiment replaces difficulty with reasoning-token counts, which yields a positive R-squared and is offered as a more LLM-aligned difficulty axis.","core_discovery":"The central claim is that high scores on human-crafted benchmarks do not necessarily reflect human-aligned capabilities in LLMs. The authors establish this by evaluating models on GAOKAO-Eval, a comprehensive, annually updated Gaokao-based benchmark with closed-book conditions and teacher-based grading, then comparing LLM scoring rates against the Rasch model's theoretical human performance curve. They find two systematic deviations: a semi difficulty-invariant scoring distribution, where the correlation between item difficulty and scoring rate is near zero, and high variance in scoring rates for items of similar difficulty. They also document grading inconsistencies among human raters for LLM responses, with an inconsistent score rate above 32% in some subjects, and recurring error patterns such as hallucinated poems or copying instead of summarizing. They further report that using o1's reasoning tokens as an alternative difficulty axis raises the Rasch fit from negative to positive, suggesting the mismatch reflects the human-aligned difficulty axis rather than only the models' deficiencies.","pith_inferences":["A direct test would be to collect human examinee scores on the same 2024 Gaokao items; if humans also deviate from the Rasch curve, the mismatch is a property of the items rather than a uniquely LLM failure.","The reasoning-token result suggests a testable hypothesis: models that spend more inference compute on harder-for-them questions will show a steeper scoring-rate curve; this could be validated across models with and without chain-of-thought.","The semi-invariance finding might partly reflect the granularity of the Elo difficulty scale or the mixture of question types; recomputing the correlations per question type could reveal whether the pattern is universal or concentrated in certain formats.","The high teacher disagreement rate raises a benchmark-design question: for subjective items, LLM answers may need a different rubric than human answers, and averaging three teachers' scores may obscure systematic oddities."],"forward_implications":["Leaderboard scores on human-crafted knowledge benchmarks should not be read as evidence of human-like reasoning, because a high aggregate score can coexist with insensitivity to item difficulty.","Benchmark designers should consider adding difficulty-response diagnostics, such as the Rasch fit and the difficulty-scoring correlation, alongside total scores.","LLM-aligned difficulty, for example based on reasoning effort, may be a more informative axis than human-perceived difficulty for evaluating and perhaps training models.","The high variance on similar-difficulty items implies reliability concerns for deploying LLMs in settings where consistent performance matters, such as tutoring or advice-giving systems."],"supporting_citations":[{"why":"Supplies the Rasch model, the theoretical human performance curve that the paper uses as the yardstick for LLM scoring patterns.","marker":"Rasch, 1993"},{"why":"Presents the Rasch model as fundamental measurement in human sciences, grounding its use as a human-performance benchmark.","marker":"Bond & Fox, 2007"},{"why":"Provides a Rasch analysis primer for practitioners, supporting the methodological choice.","marker":"Boone & Noltemeyer, 2017"},{"why":"Documents applications of Rasch measurement in educational research, used to justify the psychometric framing.","marker":"Khine, 2020"},{"why":"Gaokao-Bench, the predecessor dataset whose potential data leakage motivates GAOKAO-Eval's closed-book design.","marker":"Zhang et al., 2023"},{"why":"GAOKAO-MM, the multimodal Gaokao benchmark that GAOKAO-Eval extends and whose leakage issues it addresses.","marker":"Zong & Qiu, 2024"},{"why":"Demonstrates that larger and more instructable LLMs become less reliable, motivating the claim that high scores hide inconsistency.","marker":"Zhou et al., 2024"},{"why":"Identifies the pattern of correct answers with flawed reasoning (RM pattern), which the paper links to teacher grading inconsistencies.","marker":"Qiao et al., 2024"},{"why":"Describes automated discovery of LLM error patterns, which the paper extends to Gaokao-specific recurring mistakes.","marker":"Chen et al., 2024b"}],"fun_headline_variants":["High Gaokao scores don't mean human-like reasoning","LLM scores flat as Gaokao difficulty rises","Gaokao high scores hide weak reasoning skills","Rasch model exposes LLMs' difficulty-blind scoring"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that human examinee performance on these Gaokao items follows the Rasch curve, so that any LLM deviation from that curve is evidence of a capability mismatch; it relies on a theoretical human curve rather than collecting human scores on the same 2024 items.","fun_headline_variants_meta":{"raw":{"variants":["High Gaokao scores don't mean human-like reasoning","LLM scores flat as Gaokao difficulty rises","Gaokao high scores hide weak reasoning skills","Rasch model exposes LLMs' difficulty-blind scoring"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000449,"raw_usage":{"total_tokens":2288,"prompt_tokens":990,"completion_tokens":1298,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":1235}},"tokens_in":606,"tokens_out":1298,"duration_ms":10411,"temperature":1.0,"reasoning_tokens":1235,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:24:11.235710+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 2024 Gaokao questions on a large sample of human examinees and fit the Rasch model to their scores; if the human data also show near-zero or negative R-squared, the reported mismatch is a property of the items or the difficulty scale, not evidence about LLM capabilities.","supporting_citations":[{"cited_title":"Applying the rasch model: fundamental measurement in the human sciences, 2007","cited_arxiv_id":null,"evidence_quote":"Presents the Rasch model as fundamental measurement in human sciences, grounding its use as a human-performance benchmark."},{"cited_title":"Rasch Measurement: Applications in Quantitative Educational Research, volume 1 of Education","cited_arxiv_id":null,"evidence_quote":"Documents applications of Rasch measurement in educational research, used to justify the psychometric framing."},{"cited_title":"Gaokao-mm: A chinese human-level benchmark for multimodal models evaluation, 2024","cited_arxiv_id":null,"evidence_quote":"GAOKAO-MM, the multimodal Gaokao benchmark that GAOKAO-Eval extends and whose leakage issues it addresses."}],"review_version":1}