{"id":"a8f41d51-6b76-4133-a040-6ea4a27d0982","arxiv_id":"2608.09548","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"ELBench evaluates nine LLMs on capability, safety, teaching, and cultivation in one protocol, finding module profiles diverge, safety trades off with teaching, and all models share a style-over-fit judgment blind spot.","lead":"This paper introduces ELBench, a benchmark that evaluates large language models on four education-facing requirements at once: general capability, safety, teaching quality, and high-level judgment. It reports that top models tie on overall score while differing by module, that safety and practical teaching are anti-correlated, and that all models share a judgment blind spot favoring style over stated goals.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Appendix C's uniform-deviation analysis is computed on 'all ten evaluated models' and 5,000 responses, but the paper evaluates nine models; the cultivation blind-spot finding is not established for the stated model set.","rationale":"The stress-test pass independently examined the main text and all appendices. The most load-bearing issue is not the philosophical defensibility of the cultivation reference options, though that remains a valid concern, but a concrete internal inconsistency: Appendix C's quantitative support for the third headline finding is computed on ten models while the paper's stated evaluation set is nine. This is exactly the kind of missing-support issue the review should catch. If the tenth model is absent from Table 2, the 'common measurement protocol' claim is violated; if the ten-model counts are used, the reported statistics do not correspond to the released leaderboard. The benchmark's construction is otherwise careful: task-appropriate scoring, deterministic checks for closed-form tasks, a calibrated LLM judge with a reported human-agreement ceiling, bootstrap confidence intervals on the overall score, and explicit limitations. The reader's CONDITIONAL verdict remains appropriate; the authors should be required to audit Appendix C and either document the tenth model or recompute the cultivation statistics on the nine-model set before acceptance. The reader's weakest_assumption (single-expert references) is a different but related weakness, hence partial agreement.","tokens_in":18066,"tokens_out":12254,"duration_ms":103668,"concrete_test":"Obtain the released per-model response logs for the 500 structured-judgment items and cross-reference them against Table 2. Enumerate the actual models scored; if ten are present, identify the tenth, add it to Table 2, and rerun the overall leaderboard; if only nine are present, recompute the Appendix C statistics (140/87 correct-by-all/none, 105/84/53 common non-reference counts, and the 4,999/5,000 alignment figure) on the nine-model set. Report the new numbers and state whether the convergence (share of models on the common non-reference option around 0.97) survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5 and Table 2 state that nine models are evaluated (seven general-purpose plus InnoSpark-235B and Safe-InnoSpark). Appendix C, which provides the evidence for the third finding, repeatedly analyzes 'all ten evaluated models': 140/500 items correct by all ten, 105/84/53 items where eight, nine, or all ten models pick the same non-reference option, and 4,999 of 5,000 model responses agreeing with the deterministic grader. Since 500 structured-judgment items times ten models equals 5,000 responses, these numbers are not a typo: the uniform-deviation analysis was run on a model set different from the one in Table 2. If a tenth model exists, it is undocumented and the leaderboard may be incomplete; if not, every Appendix C statistic is wrong and the 'style-over-fit' blind spot is not supported for the stated nine models. The paper nowhere flags or explains this mismatch.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ELBench, a four-module benchmark (General Capability, Safety and Trustworthiness, Basic Education, High-Level Cultivation) for education-facing LLMs, assembled from curated public sources and a human-in-the-loop synthesis pipeline. Nine models—seven frontier general-purpose systems and two education-specialized variants—are evaluated under a common protocol, with reference/rule scoring for closed-form tasks and a calibrated LLM judge for open-ended tasks. The reported findings are: (1) the top six models are statistically tied on overall score but differ at module level, with Safety anti-correlated with Basic Education (r = -0.83); (2) Chinese-developed models lead the Safety module, mainly on region-specific normative content; and (3) education-specialized models lead neither education module, while all models share a 'style-over-fit' blind spot on the High-Level Cultivation structured-judgment task. The paper argues that module-level profiles are more informative than a single aggregate leaderboard.","tokens_in":18164,"tokens_out":11433,"duration_ms":94844,"significance":"If the issues below are resolved, ELBench would be a genuinely integrative contribution. The measurement protocol is careful in several respects: item-level bootstrap confidence intervals, paired bootstrap comparisons, a single LLM judge selected for highest agreement with a human gold set, majority voting over nine judge calls, and a uniform-deviation analysis that explicitly checks for parsing artifacts. The safety-versus-teaching correlation and the cultivation blind spot are concrete, falsifiable observations that a single aggregate leaderboard would hide. The benchmark is reusable and would give deployers a model profile rather than a rank.","major_comments":[{"comment":"Appendix C's uniform-deviation analysis is run on a ten-model set, but the paper's stated evaluation set has nine models. The structured judgment task has 500 items (Section 3.1), so ten models yield 5,000 responses and nine models yield 4,500; Appendix C reports 140 items correct by all ten, 87 correct by none, 105/84/53 items chosen by at least eight/nine/all ten models, and 4,999 of 5,000 model responses agreeing with the deterministic grader. These counts are internally consistent with ten evaluated models and are not a one-off typo. Since Section 5 invokes this analysis as the evidence for the 'style-over-fit' blind-spot finding, the finding is not established for the stated nine-model set. Please either document the tenth model and include it in all leaderboard tables, or recompute Appendix C on the nine models and confirm that the concordant-error pattern and the 4,500-response parsing check still hold.","section":"Appendix C (cf. Table 2 and Table 3)"},{"comment":"The structured judgment task is scored by exact match to a single expert reference option, and Appendix C interprets concordant selection of a non-reference option as a shared model blind spot. The paper does not report multi-expert agreement on these reference options: Section 3.2 says experts cross-review 'the correctness of reference answers' and verify in a final pass, but no quantitative inter-annotator agreement is given for the 500 structured-judgment references. If a substantial share of non-reference options are defensible, the shared 'style-over-fit' error could be a benchmark artifact rather than a model property. Please report inter-annotator agreement on the reference options with an adjudication protocol and, if agreement is not perfect, restrict the uniform-deviation analysis to items with unanimous expert references and show that the conclusion is unchanged.","section":"Section 3.2 and Appendix C"},{"comment":"The sentence 'Because both modules use tasks the models can perform, this is not a difficulty artifact' does not by itself rule out a common difficulty or grading factor; the leave-one-out recomputation is the actual supporting evidence. The statement should be rephrased so that the logical weight is placed on the stability analysis rather than on an assertion about task difficulty.","section":"Section 5, 'Safety and teaching trade off' paragraph"}],"minor_comments":[{"comment":"The claim that the top six overall scores are 'mutually indistinguishable' would benefit from an explicit statement of whether all 15 pairwise paired-bootstrap tests among the six models were conducted and whether any multiple-comparison consideration was applied.","section":"Section 5 / Appendix H"},{"comment":"Please clarify in the figure or its caption whether the safety-specialized model Safe-InnoSpark is included in the 'Chinese-developed' group, since the body text treats it separately from the four Chinese-developed general models.","section":"Figure 4"},{"comment":"The correlation r = -0.83 is computed over nine models; the leave-one-out range [-0.88, -0.79] is a stability range rather than a confidence interval, so please also report a bootstrap confidence interval for the correlation.","section":"Section 5"},{"comment":"The main text reports the selected judge's mean agreement as kappa = 0.83, while Table 8 gives 0.825; please make the rounding consistent.","section":"Appendix F, Table 8"},{"comment":"The bias taxonomy is presented without a coding procedure or inter-coder reliability; a sentence stating whether Table 6 is an informal illustration or a formal coding result would be helpful.","section":"Appendix C, Table 6"}],"recommendation":"major_revision","confidential_remarks":"The ELBench author team overlaps with the authors of EduGuardBench and ELMES (curated into the safety and Basic Education modules) and with the team behind the InnoSpark/Safe-InnoSpark models. The manuscript does not disclose this overlap. I do not read this as evidence of misconduct, but a transparency statement describing the provenance and any design choices that could affect the education-specialized models' scores would strengthen the revision. I would also ask the editor to ensure that the 'first benchmark' claim is carefully checked against OmniEduBench and SHAPE, since the related-work section distinguishes those works but the novelty claim should be precise."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know up front. This is a carefully engineered benchmark that does something genuinely new: it puts general capability, safety, basic teaching, and high-level cultivation on the same models under one protocol, and the module-level profiles surface trade-offs that a single leaderboard hides. The evaluation is also more careful than most: bootstrap CIs, paired tests, a judge selected against human gold labels with documented agreement, and a detailed parsing-artifact check. But Appendix C, which carries the 'style-over-fit' blind-spot finding, analyzes all ten evaluated models and 5,000 responses while the paper evaluates nine. That inconsistency is not flagged anywhere, and it has to be resolved before the cultivation finding is taken as established.\n\nWhat's actually new: the integrated four-module design, the newly synthesized safety and cultivation items, and the uniform-deviation analysis on the structured judgment task. The safety-teaching trade-off (r = -0.83, stable under leave-one-out) is a real and interesting empirical claim. The paper is also unusually honest about its limitations.\n\nWeak spots, in proportion. The ten-model mismatch is the big one: it could be a dropped model or an unupdated appendix, but as written every Appendix C statistic is contingent on a model set that isn't in Table 2. The cultivation finding also rests on single-expert reference options with no reported inter-annotator agreement, so the shared 'error' could partly be a benchmark artifact; the paper doesn't address that. Basic Education is only 45 items and module-level CIs aren't reported, so per-module gaps are less certain than the overall ones. The education-specialization comparison lacks a matched base model, so the headline 'education-specialized models lead neither education module' is descriptive of these two models, not a clean test of post-training. There's also some author overlap with the ELMES and EduGuardBench sources and the InnoSpark models; not a fatal issue, but worth weighing when judging independence.\n\nThis paper deserves a serious referee. The core idea and most of the evaluation are solid, and the ten-model issue is fixable. As it stands I wouldn't cite it in current form, but I'd take a careful look at a revision. It's good reading-group material for any discussion of construct validity in LLM benchmarks.","headline":"A well-built four-module education benchmark with a real contribution, but the cultivation finding rests on an unexplained ten-model analysis in a nine-model paper.","tokens_in":18776,"tokens_out":2346,"would_cite":false,"duration_ms":20608,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ELBench measures education-facing LLMs on four requirements at once and finds the modules trade off.","keywords":["ELBench","education-facing LLM evaluation","module profiling","safety and trustworthiness","basic teaching tasks","high-level cultivation","LLM-as-judge scoring","refusal behavior"],"falsifier":"Re-score the 105 items on which at least eight of ten models chose the same non-reference option, using multiple independent expert raters who are not shown the reference answer, and check whether a majority of raters also rejects the reference option; if raters split or favor the rejected option, the style-over-fit conclusion collapses.","tokens_in":17809,"feed_emoji":"🎓","tokens_out":9829,"duration_ms":77415,"temperature":0.7,"pith_summary":"ELBench is a benchmark built to test whether a large language model can be deployed in education, treating that question as four separate requirements measured under one protocol: general capability, safety and trustworthiness, basic teaching, and high-level cultivation. The paper's central claim is that no single rank can capture education-facing suitability, because on the nine models evaluated the top six overall scores are statistically tied while module leaders differ, and safety is anti-correlated with practical teaching ($r=-0.83$). It also finds that the two education-specialized models lead neither education module, and that all models share a systematic blind spot on high-level cultivation, preferring pedagogically polished answers over answers that fit the stated goal. If the benchmark is right, deployers get a way to see which requirement a model sacrifices instead of a single number, and developers get a measurable target for the missing reward signal in educational judgment.","feed_headline":"Safety and teaching quality pull AI tutors in opposite directions","feed_subtitle":"For deployers: top models tie overall; module profiles show what a single rank hides.","key_machinery":"The load-bearing instrument is the four-module evaluation protocol itself: each module is scored with a task-appropriate protocol, reference matching and deterministic rules for closed-form tasks, and rubric-based LLM judging with majority voting for open-ended tasks, with item-level bootstrap used to report confidence intervals. Within this protocol, the structured educational-judgment task, which scores 500 four-option items by exact match to a single expert reference answer, is the mechanism that exposes the shared style-over-fit blind spot. By forcing a single choice, it turns the models' common preference for gentle, elaborate, or Socratic phrasing over goal-optimal directness into a measurable, shared error pattern.","core_discovery":"On the paper's own terms, ELBench is the first benchmark to evaluate all four requirements an education-facing model must satisfy on the same models under a common protocol. The core discovery is that the four modules are not redundant: the top six models are statistically indistinguishable on overall score, yet module leaders differ substantially, and Safety and Basic Education are strongly anti-correlated ($r=-0.83$), with the correlation stable under leave-one-model-out checks. The safety module is the most discriminative, and the Chinese-developed models lead it by a margin that is largest on region-specific normative content and smaller but nonzero on universal-harm content. The two education-specialized models lead neither education module, and on the structured judgment task all models converge on the same non-reference option, favoring pedagogical style over fit to the stated cultivation goal; on 105 items at least eight of ten models pick the same non-reference option, which is why the module scores uniformly low and does not separate models.","pith_inferences":["Inference: If the safety-teaching trade-off holds beyond these nine models, benchmark designers should present a Pareto frontier between safety and teaching openness rather than a single safety-teaching score.","Inference: The uniform style-over-fit pattern may extend to other value-laden professional judgment domains, such as medical communication or social work, where fluent, empathetic phrasing can override the goal-correct response.","Inference: The cultivation module could be stress-tested by replacing single-expert references with outcome-grounded data, such as learning gains from a simulated student; if models still converge on style, the blind spot is in training data, and if they do not, the current result is a scoring artifact."],"forward_implications":["A single aggregate score is not enough for education deployment: the top six models tie overall while their module profiles disagree, so deployers should read a model's profile before choosing it.","Safety and teaching quality behave as competing objectives in this model set, so a deployment that needs both cannot be served by one compromise score.","The safety advantage of the Chinese-developed models is concentrated in region-specific normative content; universal-harm refusal gaps are smaller, making the advantage partly context-dependent.","Education-specialized models leading neither education module raises, but does not resolve, whether domain post-training keeps pace with general frontier systems.","All evaluated models share a common style-over-fit error on high-level judgment, so the benchmark identifies a training gap rather than a ranking gap on that module."],"supporting_citations":[{"why":"Supplies the curated teaching-safety and adversarial jailbreak items in the Safety module.","marker":"(Jiang et al. 2026)"},{"why":"Supplies the Basic Education scenarios and multi-turn tutoring tasks.","marker":"(Wei et al. 2025)"},{"why":"Source of contamination-resistant subject-knowledge items in General Capability.","marker":"(Wang et al. 2024b)"},{"why":"Source of Chinese-curriculum knowledge items that English-centric suites miss.","marker":"(Huang et al. 2023)"},{"why":"Source of instruction-following items.","marker":"(Zhou et al. 2023)"},{"why":"Source of the MATH-500 subset in the mathematics ladder.","marker":"(Hendrycks et al. 2021b)"},{"why":"Establishes the LLM-as-judge scoring practice used for open-ended tasks.","marker":"(Zheng et al. 2023)"},{"why":"Describes the education-specialized models evaluated in the comparison.","marker":"(Song et al. 2025)"}],"fun_headline_variants":["AI tutors: top models tie, but safety vs teaching diverges","Safety and teaching quality anti-correlate in AI tutors","ELBench: AI tutors tie overall, differ by module skills","AI tutor blind spot: all pick style over stated goal","Chinese AI models lead safety, not teaching in tutor test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The high-level cultivation result assumes each of the 500 structured judgment items has exactly one correct expert answer; if the rejected option is a defensible teaching choice, the shared 'style-over-fit' errors could be an artifact of the scoring design rather than a model blind spot.","fun_headline_variants_meta":{"raw":{"variants":["AI tutors: top models tie, but safety vs teaching diverges","Safety and teaching quality anti-correlate in AI tutors","ELBench: AI tutors tie overall, differ by module skills","AI tutor blind spot: all pick style over stated goal","Chinese AI models lead safety, not teaching in tutor test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00055,"raw_usage":{"total_tokens":2671,"prompt_tokens":1034,"completion_tokens":1637,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":650,"completion_tokens_details":{"reasoning_tokens":1553}},"tokens_in":650,"tokens_out":1637,"duration_ms":12506,"temperature":1.0,"reasoning_tokens":1553,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:17:38.087249+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-score the 105 items on which at least eight of ten models chose the same non-reference option, using multiple independent expert raters who are not shown the reference answer, and check whether a majority of raters also rejects the reference option; if raters split or favor the rejected option, the style-over-fit conclusion collapses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of Chinese-curriculum knowledge items that English-centric suites miss."},{"cited_title":"P.; Zhang, H.; Gonzalez, J","cited_arxiv_id":null,"evidence_quote":"Establishes the LLM-as-judge scoring practice used for open-ended tasks."}],"review_version":1}