{"id":"9da0ada0-bb47-4394-9399-b0a030cb9573","arxiv_id":"2504.14928","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"In simulated five-round teaching dialogues, Llama 3.1 70B Instruct produced the largest pre/post accuracy gains among 14 LLMs, and teaching effectiveness did not track model scale or benchmark reasoning scores.","lead":"The paper introduces EducationQ, a multi-agent testbed where a simulated student takes a quiz, talks with a language model teacher for five rounds, and retakes the same questions. It reports that smaller open models such as Llama 3.1 70B can produce larger learning gains than bigger commercial models, and that general reasoning skill does not predict teaching skill.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No-teacher control missing: ALG may measure re-exposure and hint-following, not teaching, because the post-test reuses the same questions and includes the full dialogue.","rationale":"The paper's central claim is an ordering of ALG values from Table 4, and ALG is defined by Eq. 1 as the difference between post-test and pre-test accuracy. Section 4.7 states that the post-test prompt includes the student's pre-test reasoning and the full teacher-student dialogue, and the post-test uses the same questions as the pre-test. No control condition exists in which the student retakes the test without a teacher or with a content-free teacher. Therefore, a positive ALG can be produced by re-reading the question and one's own prior reasoning, by following answer-relevant hints in the dialogue, or by actually learning. The case studies in F.1 and F.2 show teacher dialogues converging on the correct option and post-test answers echoing those hints, so hint-following is plainly present. Because teacher models differ in how much answer-relevant content they emit, the Table 4 ranking may reflect context-conditioned answering rather than teaching quality. The 78% human-expert agreement validates the evaluator agent's comparative quality judgments on 50 UIC pairs but does not establish that ALG measures durable learning. This is the load-bearing concern, and it is the same weakest assumption the reader identified. The reader's conditional verdict is appropriate; the proposed no-teacher and generic-teacher control would directly settle whether the confound changes the ranking.","tokens_in":33538,"tokens_out":4300,"duration_ms":41796,"concrete_test":"Run a no-teacher and generic-teacher control on a stratified 500-question subset of the 1,498 questions using the same Llama 3.1 70B Instruct student agent. Condition A: post-test immediately after pre-test with the same question and pre-test reasoning but no dialogue. Condition B: five rounds of a fixed, content-free teacher response such as 'Please think again about the question' followed by the same post-test. Compute baseline ALG for both conditions, then subtract it from each teacher's ALG in Table 4 and apply a paired significance test. If the adjusted ranking changes, e.g., Llama 3.1 70B no longer significantly outperforms Gemini 1.5 Pro, the reported non-linear scaling result is an artifact of re-exposure or hint-following rather than teaching effectiveness.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is an ordering of ALG values in Table 4, but ALG (Eq. 1) is computed from a post-test that, per Section 4.7, reuses the same questions and includes both the student's pre-test reasoning and the full five-round teacher-student dialogue. There is no no-teacher control condition, so a positive ALG can arise from re-reading the question, recognizing answer-relevant hints in the dialogue, or genuinely learning. The case studies in F.1 and F.2 show teacher dialogues converging on the correct option and the post-test answers echoing those hints, demonstrating that hint-following is present. Because teacher models differ in how much answer-relevant content they emit, the Table 4 ranking may track context-conditioned answering rather than teaching quality. The 78% human-expert agreement validates evaluator-agent judgments on 50 comparative pairs, but it does not establish that ALG measures durable learning. This confound is not acknowledged in the Limitations section, and it directly undermines the headline claim that smaller open-source models outperform larger commercial ones in teaching contexts.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces EducationQ, a multi-agent dialogue framework for evaluating LLMs' teaching capabilities. The framework pairs a fixed student agent (Llama 3.1 70B Instruct) with a teacher agent under evaluation across five dialogue rounds per question, using a pre-test/post-test design on 1,498 questions from GPQA Diamond and a newly constructed MMLU-Pro Stratified set. The main outcome metric is Absolute Learning Gain (ALG), defined as the difference between post-test and pre-test accuracy. Testing 14 LLMs, the authors report that Llama 3.1 70B Instruct achieves the highest ALG (11.01%), followed by Gemini 1.5 Pro 002 (7.48%), while larger commercial models such as GPT-4o-mini score much lower (2.44%). They supplement this with qualitative evaluator-agent analyses, expert case studies, and a human-alignment study reporting 78% agreement. The central claim is that teaching effectiveness does not correlate linearly with model scale or general reasoning ability.","tokens_in":33756,"tokens_out":3690,"duration_ms":33692,"significance":"If the measurement of ALG were robust, this would be a valuable contribution: it provides a large-scale, reproducible, multi-agent methodology for an important but under-benchmarked capability (LLM-as-teacher), with a substantial released corpus of teacher-student dialogues and an ablation study across student models. The external human-expert alignment check (78% on 50 pairs) is a genuine strength, as are the open-source commitment and the explicit content-boundary controls preventing direct answer disclosure. However, the main quantitative claim rests on ALG in a pre/post design with identical questions and a post-test prompt that includes the full teaching dialogue; without a no-teacher control, the framework cannot separate genuine learning from re-exposure effects and hint-following. The ranking of models in Table 4 is therefore not yet established as a ranking of teaching quality. The finding that small open models can beat large commercial ones is interesting and plausible, but the current evidence does not yet support it at the level of statistical confidence required for the paper's headline claim.","major_comments":[{"comment":"The post-test protocol reuses the exact same questions as the pre-test and, per Section 4.7, includes both the student's pre-test reasoning and the full teacher-student dialogue in the prompt. Consequently, a positive ALG can result from re-exposure to the question text or from answer-relevant cues embedded in the dialogue, rather than from durable learning. The framework has no no-teacher control condition (e.g., a post-test with the same question but no teaching dialogue), so the ranking in Table 4 may reflect how well each teacher happens to steer the student toward the correct option within the dialogue context. The case studies in Appendix F.1 and F.2 illustrate this: the teacher dialogues converge on the correct option, and the post-test answers closely echo the reasoning scaffolded by the teacher. This is not acknowledged in the Limitations section. I recommend adding a no-teacher control (post-test with only the question and the student's own pre-test reasoning) and reporting ALG relative to that baseline, ideally accompanied by a retention test on non-overlapping questions to assess durable learning.","section":"Section 4.7 and Eq. (1)"},{"comment":"The overall ALG values in Table 4 are reported as point estimates with no confidence intervals, standard errors, or significance tests. Several adjacent rankings differ by fractions of a percentage point (e.g., Hermes 3 Llama 3.1 70B at 4.14%, Mistral Nemo at 3.94%, Claude 3.5 Sonnet at 3.81%). The stability study in Table 3 shows run-to-run variance for three models on GPQA-main, but these variances are not applied to the main results; for example, the reported variance of 0.01246 for Llama 3.1 405B corresponds to a standard deviation of about 0.11 percentage points, which is comparable to some of the gaps in Table 4. Without uncertainty quantification or pairwise significance tests, the specific ordering of models in Table 4 is not statistically supported. Please provide bootstrap confidence intervals over questions for each ALG and report which pairwise differences are significant after appropriate multiple-comparison correction, or explicitly state which differences are within noise.","section":"Table 4 and Section 7.1"},{"comment":"The 78% human-expert agreement validates the evaluator agent's comparative judgments of teaching behaviors on 50 anonymized dialogue pairs, which is a useful external check. However, it does not validate that ALG measures durable learning. The human experts were asked to rate teaching behaviors and to check for direct answer disclosure; they were not asked whether the student's post-test performance reflects understanding that would transfer to new questions. Given that the post-test prompt includes the dialogue content (Section 4.7), the experts' confirmation of 'no direct answer disclosure' does not rule out that students are following embedded hints rather than genuinely learning. Please either add a transfer test (post-test on previously unseen but related questions) or explicitly scope the paper's claims from 'teaching quality' to 'context-conditioned answering in the presence of a dialogue.'","section":"Section 8.2"}],"minor_comments":[{"comment":"The word 'fudamental' in the Conclusion should be 'fundamental.'","section":"Conclusion"},{"comment":"The phrase 'protential impacts' should be 'potential impacts.'","section":"Model Content Limitations"},{"comment":"The heading 'Dimensons' is misspelled; it should be 'Dimensions.'","section":"Appendix A.3.3"},{"comment":"The reference to 'Rob Wass and Clinton Golding and. 2014' contains a stray 'and' and should be cleaned up.","section":"References"},{"comment":"The case study in Table 22 uses inconsistent teacher labels: the header refers to Teacher 1 and Teacher 2, while the evaluator analysis refers to Teacher A and Teacher B, making it difficult to map the verdicts back to the models. Please align the labeling throughout.","section":"Appendix F.1"},{"comment":"The student model is referred to as 'Mistral Nemo 12b' in Section 4.1 but as 'Mistral Nemo' in other tables; please standardize the model naming.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is promising but the central empirical claim is currently undercut by a missing control condition and a lack of statistical inference. I would encourage the editor to treat the revision seriously: the authors need to add a no-teacher control and report uncertainty on the rankings. The qualitative and human-alignment components are solid and should be preserved. There is also a minor concern that the paper's framing of 'teaching effectiveness' may overstate what the ALG metric measures even after the control is added, since the post-test reuses the same questions; the authors should recalibrate their language accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper builds the most complete automated setup I have seen for evaluating LLMs as teachers: a fixed student agent, a teacher agent under test, and an evaluator agent, with pre/post accuracy as the headline. The genuinely new pieces are the fixed-student formative assessment design, the MMLU-Pro Stratified dataset, and a ranking of 14 models that does not track raw scale or general reasoning. It also ships code and detailed appendices, and the 78% human agreement on qualitative pairwise verdicts is real evidence that the evaluator agent can identify effective teaching behaviors. Credit where due: the framework is reproducible, and the qualitative case studies are useful teaching examples in their own right.\n\nThe soft spot is exactly what the stress-test note says. ALG is post-minus-pre accuracy, but the post-test reuses the same questions and the student receives the full five-turn dialogue plus its own pre-test reasoning. There is no no-teacher control. So a positive gain can come simply from re-exposure to a previously seen question, or from recognizing answer-relevant hints in the dialogue, rather than from durable learning. The paper's own case studies show this dynamic: the dialogues converge on the correct option, and the post-test answers echo those hints. Since teacher models differ in how much answer-relevant content they emit, the Table 4 ranking may partly measure context-conditioned answering, not teaching quality. This confound is not acknowledged in the Limitations section, and that is a real omission.\n\nThis is a load-bearing flaw for the headline claim that smaller open-source models outperform larger commercial ones. It does not kill the framework; it kills the current empirical claim. The fix is straightforward and in-scope: run a no-teacher control condition where the student takes the post-test without any dialogue, and report confidence intervals or significance tests on the ALG differences, because several adjacent rankings differ by only a few tenths of a point. The student-model ablation is helpful, but it does not address re-exposure or hint-following.\n\nIs this worth a serious referee? Yes. The methodology is novel and the qualitative-plus-alignment component is worth building on, but a referee should insist on a control before accepting the scale-non-correlation finding. \n\nRecommendation: send it to peer review, expect heavy revision. Read it for the framework and the dataset, not for the ranking.\n\nBest,\n[Your name]","headline":"A genuinely new framework for evaluating LLMs as teachers, with a central empirical claim that is not yet proven because the post-test lacks a no-teacher control; still worth refereeing, with heavy revision expected.","tokens_in":34279,"tokens_out":1779,"would_cite":true,"duration_ms":18317,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper reports that LLM teaching ability is a distinct, measurable capability that does not track model scale or reasoning benchmarks, with a 70B open model outperforming larger commercial models in simulated student dialogues.","keywords":["LLM evaluation","teaching capability","multi-agent dialogue","formative assessment","learning gain","LLM-as-teacher","GPQA","MMLU-Pro"],"falsifier":"Run the identical pre/post protocol with a no-teacher control and with a post-test made of new but matched questions; if the no-teacher control yields a similar ALG, or if gains disappear on held-out questions, the EducationQ ranking is measuring test familiarity rather than teaching.","tokens_in":33332,"feed_emoji":"🎓","tokens_out":7223,"duration_ms":59417,"temperature":0.7,"pith_summary":"EducationQ proposes a way to test how well large language models teach, separate from how much they know. It stages a five-round dialogue between a fixed simulated student and the model being evaluated, on 1,498 graduate-level questions, and scores the teacher by the student's accuracy gain from pre-test to post-test. The paper reports that this teaching ranking does not follow model scale or general reasoning scores: Llama 3.1 70B Instruct tops the list with an average 11.01% gain, ahead of Gemini 1.5 Pro 002 at 7.48%, Llama 3.1 405B Instruct at 6.14%, and GPT-4o-mini at 2.44%. Expert teachers agreed with the automated evaluator's choice in 78% of anonymised comparisons, and the authors use this to argue the qualitative assessment is meaningful. If the finding holds, selecting and building educational LLMs requires measuring interactive pedagogy, not just benchmark accuracy.","feed_headline":"A 70B open model out-teaches larger rivals in teaching test","feed_subtitle":"In simulated student dialogues, Llama 3.1 70B gains 11.0% accuracy vs 7.5% for Gemini 1.5 Pro and 2.4% for GPT-4o-mini.","key_machinery":"The load-bearing mechanism is the multi-agent formative-assessment loop: a teacher agent, a fixed student agent (Llama 3.1 70B Instruct), and a GPT-4o evaluator agent. The teacher is given the student's pre-test reasoning and correctness, but never the answer options, and conducts five rounds of questioning; between pre-test and post-test the same questions are readministered. Teaching effectiveness is the ALG metric, $\\mathrm{ALG} = \\mathrm{ACC}_{\\mathrm{post}} - \\mathrm{ACC}_{\\mathrm{pre}}$; stability and uniqueness metrics (PNIR, CSS, UIC) support the ranking, and a 17-dimension evaluator rubric turns dialogues into qualitative scores that correlate with human judgments. The critical design choice is the enforced information boundary: teachers cannot see answer options, so gains are attributed to pedagogical dialogue rather than answer leakage.","core_discovery":"The paper's central claim is that teaching is a measurable, separable capability in LLMs, and that its ranking cannot be inferred from the usual benchmarks. EducationQ's triadic setup—teacher, student, evaluator—produces a quantitative teaching score per model: Absolute Learning Gain (ALG), the percentage-point improvement in a fixed student agent's accuracy from pre-test to post-test after five teacher turns. On 1,498 questions spanning 13 disciplines and 10 difficulty levels, the best teacher is Llama 3.1 70B Instruct (ALG 11.01%), followed by Gemini 1.5 Pro 002 (7.48%), with Llama 3.1 405B Instruct at 6.14%, OpenAI o1-mini at 5.84%, and GPT-4o-mini at 2.44%. The authors report model-specific teaching styles—progressive questioning and scaffolding for Llama 3.1 70B, targeted feedback for Gemini 1.5 Pro 002, and reasoning-heavy support for o1-mini—and note that no expert reviewer observed a teacher revealing answers. They take these results to challenge the assumption that larger scale or higher general intelligence directly improves teaching.","pith_inferences":["The paper has no no-teacher control: the same gain could partly come from re-reading the question or following hints embedded in the dialogue, so a no-teacher and a hint-only baseline would reveal how much of the 11% is teaching.","Because the post-test reuses the identical pre-test questions and includes the dialogue in the student's context, the measured gains may reflect short-term answer reshaping rather than durable understanding; a held-out, same-topic post-test would distinguish these.","Since the student is a single 70B model, the ranking may be specific to that simulated learner; using novice or grade-school student personas could reorder the teachers.","An extension the authors do not pursue is to use the same evaluator loop to give formative feedback to teacher models, turning the benchmark into a training signal."],"forward_implications":["Educational model selection should treat teaching ability as an independent axis, not as a corollary of reasoning benchmarks.","Dialogue-based simulated evaluation can replace some human panels for ranking teaching quality, with the 78% agreement supporting scaled qualitative review.","Models have complementary teaching styles, so a tutoring system could route students to the teacher model best matched to the subject and the student's state.","Scaling parameter count alone is not a route to better teaching; optimising questioning and feedback behaviour is."],"supporting_citations":[{"why":"Source of the GPQA and GPQA Diamond questions used in the evaluation dataset.","marker":"Rein et al., 2023"},{"why":"Source of MMLU-Pro and basis for the difficulty-stratified MMLU-Pro Stratified subset.","marker":"Wang et al., 2024b"},{"why":"Provides the definition of formative assessment that the dialogue protocol operationalizes.","marker":"William, 2011"},{"why":"Supplies the Zone of Proximal Development scaffolding principle the authors say top teachers apply.","marker":"Vygotsky, 1978"},{"why":"Supports using GPT-4 as a human-aligned evaluator judge in the qualitative analysis.","marker":"Zheng et al., 2023"},{"why":"Grounds the learning-gain measurement used as the primary quantitative outcome.","marker":"McGrath et al., 2015"},{"why":"Models the informal formative assessment teacher-student dialogue on which EducationQ is based.","marker":"Sezen-Barrie and Kelly, 2017"}],"fun_headline_variants":["Smaller open LLMs outperform giants in teaching test","Llama 3.1 70B tops GPT-4o-mini in simulated teaching","Teaching skill doesn't scale with LLM size","Why a 70B model beats 405B as a teacher"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the post-test gain is caused by the teacher's dialogue, but since the post-test repeats the same questions and includes the dialogue in the student's context, part of the gain could come from re-exposure or hint-following rather than durable teaching.","fun_headline_variants_meta":{"raw":{"variants":["Smaller open LLMs outperform giants in teaching test","Llama 3.1 70B tops GPT-4o-mini in simulated teaching","Teaching skill doesn't scale with LLM size","Why a 70B model beats 405B as a teacher"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000199,"raw_usage":{"total_tokens":1421,"prompt_tokens":1042,"completion_tokens":379,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":658,"completion_tokens_details":{"reasoning_tokens":305}},"tokens_in":658,"tokens_out":379,"duration_ms":3845,"temperature":1.0,"reasoning_tokens":305,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:36:48.894068+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical pre/post protocol with a no-teacher control and with a post-test made of new but matched questions; if the no-teacher control yields a similar ALG, or if gains disappear on held-out questions, the EducationQ ranking is measuring test familiarity rather than teaching.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the definition of formative assessment that the dialogue protocol operationalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds the learning-gain measurement used as the primary quantitative outcome."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Models the informal formative assessment teacher-student dialogue on which EducationQ is based."}],"review_version":1}