{"id":"62584ede-011b-4756-86af-9540404e9d5e","arxiv_id":"2505.00057","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Eight LLM APIs score 10% to 64% accuracy on a hand-picked Gaokao math set, with rankings that contradict the paper's own weighted scores.","lead":"This report runs eight large language model APIs on selected Gaokao high school math questions and compares their accuracy, response time, logical reasoning, and creativity. It is an example of a routine model benchmark, but the paper's internal numbers are inconsistent and no data or code is released.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Internal scoring contradictions invalidate the central ranking: Table 15's weighted scores cannot be monotonically sorted into the claimed order, and Table 17's accuracy ranking contradicts Table 12.","rationale":"The reader's REJECT verdict is correct, but my primary load-bearing concern is internal inconsistency in the reported ranking rather than the (also serious) unvalidated grading pipeline. The weighted scores in Table 15 cannot be sorted into the stated ranking under either a higher-is-better or lower-is-better reading, so the conclusion of Section 5.5.3 is not reproducible from the paper's own tables. Table 17's accuracy claims are likewise inconsistent with Table 12. These are not presentation issues: the contribution is the ranking and the accuracy numbers, and the contradictions are directly about those numbers. The grader concern in Section 4.1.1 is corroborating and still important: without a validated gold standard, the accuracy values themselves are not established, and the observed pattern (multiple-choice accuracy below random, comprehensive accuracy near ceiling) is a warning sign. Because the central claim fails on internal evidence and the methodology is under-specified, there is no reason to adjust the reader's REJECT. I mark agreement as partial because the reader's weakest_assumption emphasizes the grader, while I see the self-contradictory scoring as the decisive independent failure.","tokens_in":15534,"tokens_out":9468,"duration_ms":84646,"concrete_test":"Recover the exact normalization and weighting formula behind Table 15 (from the authors or code) and recompute the eight weighted scores from Table 14; then sort them and compare to the claimed ranking. If no monotone reading of the scores (higher-is-better or lower-is-better) yields Qwen > GLM > Hunyuan > Spark > ERNIE > gemma > Meta > Yi, the central ranking is unsupported by the report's own data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central deliverable is a ranking of eight LLMs from measured metrics. Section 5.5.3 and Table 15 give weighted scores: Qwen2.5-7B-Instruct 0.8124917, GLM-4-Flash 0.983135, hunyuan-lite 0.7870671, Spark-lite 0.769745, ERNIE-Speed-128K 0.768396, gemma-2-9b-it 0.7422361, Meta-Llama-3.1-8B-Instruct 0.483141, Yi-34B 0.6072355. The claimed ranking is Qwen, GLM, Hunyuan, Spark, ERNIE, gemma, Meta, Yi. Sorting the scores either ascending or descending gives neither that order nor a consistent monotone mapping: Qwen is below GLM, and Yi is below Meta. The final ranking is therefore not derivable from the report's own score table. Independently, Table 17 states Spark-lite and Qwen rank first in accuracy, while Table 12 gives ERNIE-Speed-128K the highest comprehensive-question accuracy (0.95) and Qwen higher multiple-choice and fill-in-the-blank accuracies than Spark-lite; no aggregation rule is given that makes Spark-lite the accuracy leader. Section 4.1.1 compounds the problem: multiple-choice scoring is exact string matching, fill-in-the-blank is an unspecified fuzzy method, and comprehensive questions use an unvalidated AI score with only a 10% expert sample. The multiple-choice accuracies of 0.10–0.20 sit below random for all models while comprehensive accuracies are 0.70–0.95, a pattern consistent with grader failure rather than model ability. Since the ranking depends entirely on these numbers, the central claim fails on the paper's own evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports an empirical evaluation of eight large language models (GLM-4-Flash, ERNIE-Speed-128K, Spark-lite, hunyuan-lite, Qwen2.5-7B-Instruct, gemma-2-9b-it, Meta-Llama-3.1-8B-Instruct, and Yi-34B) on a set of Chinese Gaokao mathematics questions from 2019 to 2023. The authors define metrics for accuracy, response time, logical reasoning, guidance, and creativity, then present per-model tables and figures, and conclude with a comprehensive ranking in which Qwen2.5-7B-Instruct is ranked first. The central deliverable is this ranking and the claim that the data provide a solid quantitative foundation for assessing LLMs in education.","tokens_in":15905,"tokens_out":2575,"duration_ms":25803,"significance":"If the evaluation pipeline and ranking were internally consistent, this study could provide a useful practical benchmark for educational deployment of LLMs, because it uses real exam questions and compares eight readily available models across multiple dimensions. The paper also has the strength of separating single-solution and multi-solution responses, and of reporting response-time statistics. However, the significance is undermined by multiple internal contradictions in the reported numbers, and by an unvalidated automatic grading pipeline. As presented, the results do not support the paper's central claims, so the contribution is currently not usable by the community.","major_comments":[{"comment":"The weighted scores in Table 15 are not consistent with the stated ranking. Qwen2.5-7B-Instruct is given a weighted score of 0.8124917 but is ranked first, while GLM-4-Flash has a score of 0.983135 and is ranked second. Similarly, Yi-34B has a score of 0.6072355 and is ranked eighth, whereas Meta-Llama-3.1-8B-Instruct has a lower score of 0.483141 and is ranked seventh. A higher weighted score should correspond to a better rank, but this ordering is reversed in two places. Because the ranking is the main conclusion of the paper, this contradiction is load-bearing.","section":"Section 5.5.3, Table 15"},{"comment":"Table 17 states that Spark-lite and Qwen2.5-7B-Instruct rank first in accuracy, but Table 12 shows ERNIE-Speed-128K with the highest comprehensive-question accuracy (0.95) and Spark-lite with the lowest multiple-choice accuracy (0.10). No aggregation rule is provided that would make both Qwen and Spark-lite the accuracy leaders, and the stated accuracy ranking is therefore not reproducible from the paper's own data. This inconsistency directly affects the per-metric conclusions in Section 5.6.","section":"Section 5.6, Tables 16 and 17 versus Section 5.2.1, Table 12"},{"comment":"The automatic grading pipeline is not validated against an independent human-labeled gold set, and the resulting numbers are implausible in a way that suggests grader failure. Multiple-choice accuracies of 0.10 to 0.20 for all models are at or below random guessing for four-option questions, while the same models show comprehensive-question accuracies of 0.70 to 0.95. The paper does not explain how a model could score below chance on multiple-choice questions while scoring very highly on open-ended comprehensive questions. Since the accuracy metric is central to the ranking, the lack of grader validation is a load-bearing weakness.","section":"Section 4.1.1 and Tables 12 and 17"},{"comment":"The creativity composite weighted score used for the final ranking is not fully specified. The paper lists five dimensions (correctness, solution diversity, solution richness, time taken, and solution complexity) and shows a bar chart labeled 'Factor percentage', but it never gives the exact formula or the weight vector that maps the raw dimensions to the weighted scores in Table 15. Without this information, the top-level ranking in Table 15 is not reproducible, and the claim that one model 'performs the best when considering all metrics and weights' cannot be independently verified.","section":"Section 4.5 and Figure 24"}],"minor_comments":[{"comment":"The definitions for 'Advanced Questions' and 'Difficult Questions' are duplicated verbatim, and both contain the same error about difficulty factors below 0.3, which makes the table confusing.","section":"Table 3"},{"comment":"Both tables are labeled 'Data synthesis tables', and one of them contains a column headed with the Chinese word for 'difficult' rather than an English description; the duplicated numbering should be corrected.","section":"Tables 10 and 11"},{"comment":"Several figure captions contain typos, such as 'Average orrectness' instead of 'Average correctness', and the axis labels in Figures 14 and 15 are similarly misspelled.","section":"Figures 12-16"},{"comment":"The formula for the comprehensive score includes the term '(Guidance Attempts /1 × 0.6)', where the division by 1 is meaningless and likely a typographical error; the intended formula should be clarified.","section":"Section 4.3"},{"comment":"The model name 'ERNIE-Speed-128K' is misspelled as 'LRNE-Speed-128K' in the text of Section 5.1.3, which is distracting and should be corrected.","section":"Section 5.1.3"}],"recommendation":"reject","confidential_remarks":"The manuscript appears to be an evaluation report rather than a peer-reviewed research paper, and the internal numerical contradictions suggest that the results have not been carefully checked. In particular, the inconsistency between the weighted scores and the final ranking in Section 5.5.3, and the contradiction between Tables 12 and 17 regarding the accuracy leader, cannot be resolved by minor edits. Even if the ranking were corrected, the unvalidated grading pipeline would require a complete re-evaluation with a human-labeled gold set. I recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's my read on 2505.00057.\n\nThe paper tries to do something useful: compare eight LLM APIs on Gaokao mathematics questions across accuracy, latency, logic, guidance, and creativity, with two prompt conditions. That design is reasonable for a practical evaluation. But the execution doesn't hold up. The central deliverable — the ranking in Section 5.5.3 — is contradicted by the paper's own Table 15. Qwen2.5-7B-Instruct is declared first with weighted score 0.812, while GLM-4-Flash scores 0.983 and is placed second. Yi-34B scores 0.607 and is ranked eighth while Meta-Llama-3.1-8B-Instruct scores 0.483 and is ranked seventh. No monotone rule maps those numbers to that order. Table 17 also states Spark-lite ranks first in accuracy, but Tables 12 and 13 show ERNIE-Speed-128K with the highest comprehensive-question accuracy (0.95) and Qwen with higher multiple-choice and fill-in-the-blank accuracy than Spark-lite. There's no aggregation that makes Spark the accuracy leader.\n\nThe soft spots are load-bearing, not cosmetic. Section 4.1.1's grader is unvalidated: exact string matching for multiple-choice, an unspecified fuzzy method for fill-in-the-blank, and an AI-assisted score with only a 10% expert sample for comprehensive questions. No independent gold set is used. The reported multiple-choice accuracies (0.10–0.20) are below random for four-choice items, while comprehensive accuracy is 0.70–0.95; that pattern strongly suggests the grader is broken, not that models are worse at multiple choice.\n\nThe presentation also undermines confidence. Table 9's prompt definitions are swapped: the single-solution prompt asks for multiple solutions, and the multi-solution prompt asks for only the answer. The weighted creativity score uses arbitrary weights, thresholds like the 80% multi-solution correctness rule and the adaptive fuzzy threshold are unsupported, and no error bars, statistical tests, or data/code are provided. The reference list is thin and doesn't cite existing Gaokao benchmarks (e.g., GaokaoBench), so the paper situates itself poorly in the literature.\n\nCredit where it's due: the question selection across exam years and difficulty bins is thoughtful, and the idea of logging multi-solution richness and complexity is a reasonable direction for evaluating creativity. But those seeds aren't enough to rescue an evaluation whose headline results are internally inconsistent.\n\nWho is this for? Maybe a practitioner who wants a scrap of comparative latency data across these APIs — but even that is uncertain given no release of the prompts or data. This paper should not be sent to an external referee; it needs major methodological work before it merits review.\n\nMy advice: desk-reject, and tell the authors to fix the table inconsistencies, validate the grader, and release data if they resubmit.","headline":"The paper's own tables sink its ranking: weighted scores contradict the claimed order, accuracy claims switch leaders, and the grading pipeline is unvalidated.","tokens_in":16449,"tokens_out":3980,"would_cite":false,"duration_ms":36251,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper ranks eight LLMs on Gaokao math and crowns Qwen2.5-7B-Instruct, even as multiple-choice accuracy sits at 10-20 percent.","keywords":["large language models","Gaokao mathematics","LLM evaluation","automatic grading","chain-of-thought prompting","educational technology","response time benchmarking","creativity evaluation"],"falsifier":"Take a random sample of the model outputs previously graded by the pipeline, regrade them with independent human experts using the standard answers, and measure agreement; if agreement on multiple-choice or comprehensive questions is low or systematically tilted in one direction, the reported accuracy differences and rankings will not survive.","tokens_in":15289,"feed_emoji":"🧮","tokens_out":8311,"duration_ms":77119,"temperature":0.7,"pith_summary":"This paper tries to establish a quantitative benchmark for large language models in education by testing eight LLM APIs on Chinese college-entrance (Gaokao) mathematics questions from 2019 to 2023. The authors categorize questions by type and difficulty, collect single- and multi-solution answers from each model, and score them on accuracy, response time, logical reasoning, prompt guidance, and creativity. Their main conclusion is that model performance depends heavily on question format: multiple-choice accuracy is strikingly low (0.10-0.20 for every model), while comprehensive-question accuracy reaches 0.70-0.95; in the overall weighted ranking they name Qwen2.5-7B-Instruct first and GLM-4-Flash second, though their own Table 15 gives GLM-4-Flash a higher weighted score. If the evaluation pipeline is sound, the results would give educators and developers a usable basis for choosing models and for steering them with prompts.","feed_headline":"LLMs get 10-20% on Gaokao multiple choice, up to 95% on open math","feed_subtitle":"Report scores eight models on speed, logic and creativity to guide tutoring choices; accuracy varies hugely by question type.","key_machinery":"The carrying instrument is the paper's four-way evaluation pipeline. For multiple-choice questions it uses direct string matching after answer normalization; for fill-in-the-blank it uses fuzzy semantic matching built on word-vector embeddings; for comprehensive questions it uses an AI-assisted neural grader plus random expert re-scoring of one tenth of the answers; and creativity is measured by decomposing multi-solution outputs with regular expressions and greedy scanning, then scoring correctness, richness, complexity, logic, and time. These scores feed weighted tables, especially Table 15, that produce the final model ranking.","core_discovery":"The report's central discovery, stated on its own terms, is a comparative performance profile of eight LLMs on a curated set of Gaokao math problems. It reports that all models answer only 10-20% of multiple-choice questions correctly under direct string matching, whereas on comprehensive questions they score between 0.70 and 0.95, with ERNIE-Speed-128K at 0.95 and Qwen2.5-7B-Instruct at 0.85; multi-solution accuracy is highest for Qwen2.5-7B-Instruct (0.64) and ERNIE-Speed-128K (0.63). It also reports that prompt framing as a mathematician, exam student, or teacher improves accuracy by up to 22% (GLM-4-Flash under the student prompt). The intended headline conclusion is that Qwen2.5-7B-Instruct performs best across all metrics and weights; the table accompanying that conclusion assigns GLM-4-Flash the higher weighted score (0.983 vs 0.812), so the ranking as printed is internally inconsistent.","pith_inferences":["The 10-20% multiple-choice accuracy is likely an artifact of strict string matching plus answer-format mismatch; if the grader were validated against human labels, the choice accuracy would probably rise, but the paper does not test this.","Because Table 15 gives GLM-4-Flash (0.983) a higher weighted score than the prose winner Qwen2.5-7B-Instruct (0.812), the printed ranking cannot be reproduced from the paper's own table; any user should recompute the weights before trusting the winner.","The described pipeline could be turned into a proper benchmark by human double-scoring a random sample of outputs; until then the accuracy numbers and rankings are conditional on the unvalidated automatic grader.","Neighboring educational uses, such as personalized tutoring, automatic homework grading, and exam generation, inherit all of these caveats, so the report is best read as a pilot rather than a foundation."],"forward_implications":["If the accuracy figures hold, LLMs in their current form are unreliable for standardized multiple-choice math assessment, scoring at or below chance, while being much more useful on open-ended problems.","Prompt guidance is a practical lever: framing the task with an expert or student persona raised accuracy for most models by 2-22 percentage points, so deployment should tune prompts, not just models.","Speed separates models by an order of magnitude (0.32 s vs 4.50 s on multiple-choice; 688 s vs 7400 s for multi-solution batches), which matters for real-time tutoring applications.","The weighted score is intended as a composite education-suitability score, giving later work a starting point even though the printed ranking is internally inconsistent."],"supporting_citations":[{"why":"Supplies the CCoT method used to limit output length when scoring logical reasoning.","marker":"1"},{"why":"Provides the word-vector representations that the fuzzy matching step is built on.","marker":"2"},{"why":"Supports the multi-dimensional text-matching method used for fill-in-the-blank grading.","marker":"3"},{"why":"Supplies dependency parsing used by the AI-assisted grader for comprehensive questions.","marker":"4"},{"why":"Provides regular expressions that decompose multi-solution outputs into separate solutions.","marker":"5"},{"why":"Provides the greedy pattern used to scan the decomposed solution list.","marker":"6"},{"why":"Defines the 95th-percentile response-time benchmark used for speed scoring.","marker":"7"},{"why":"Connects chain-of-thought prompting to LLM deployment, underpinning the guidance metric.","marker":"8"}],"fun_headline_variants":["Eight LLMs flunk Gaokao multiple choice, ace open math","Gaokao math: LLMs score 10-20% on MCQs, up to 95% open","LLMs fail MCQs but shine on open Gaokao math","Prompt as student lifts LLM Gaokao scores by 22%","Inconsistent ranking in LLM Gaokao test: Qwen vs GLM"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire ranking assumes the automatic grading pipeline labels model outputs correctly, even though the paper never checks that pipeline against an independent human-scored gold set.","fun_headline_variants_meta":{"raw":{"variants":["Eight LLMs flunk Gaokao multiple choice, ace open math","Gaokao math: LLMs score 10-20% on MCQs, up to 95% open","LLMs fail MCQs but shine on open Gaokao math","Prompt as student lifts LLM Gaokao scores by 22%","Inconsistent ranking in LLM Gaokao test: Qwen vs GLM"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000806,"raw_usage":{"total_tokens":3540,"prompt_tokens":948,"completion_tokens":2592,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":2486}},"tokens_in":564,"tokens_out":2592,"duration_ms":18361,"temperature":1.0,"reasoning_tokens":2486,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:59:34.944332+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the model outputs previously graded by the pipeline, regrade them with independent human experts using the standard answers, and measure agreement; if agreement on multiple-choice or comprehensive questions is low or systematically tilted in one direction, the reported accuracy differences and rankings will not survive.","supporting_citations":[],"review_version":1}