{"id":"b1764e82-bc34-45b4-a63f-dd572c2d7292","arxiv_id":"2506.18710","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors release an open benchmark of 1,143 pedagogical knowledge questions from Chilean teacher exams and report accuracy, cost, and size trade-offs for 97 large language models.","lead":"This paper introduces a new benchmark with 1,143 multiple-choice questions drawn from Chilean teacher exams to test how well large language models understand pedagogy, not just subject content. It reports scores for 97 models, showing a wide range from 28% to 89% accuracy, and analyzes the trade-off between accuracy and cost.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Filtering of ~10% of MCQs based on LLM performance may inflate scores and bias model comparisons; independent validation of the excluded questions is needed.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the post-hoc exclusion of ~10% of questions based on LLM performance. This is the single most serious threat to the central claim that the benchmark provides a valid measure of pedagogical knowledge. The paper is otherwise strong: it provides open data and code, evaluates 97 models, reports bootstrap confidence intervals, checks answer-position bias, and discusses limitations openly. However, the filtering step is not adequately validated. The authors say manual inspection found flaws, but they do not provide details on the number of excluded questions, the nature of flaws, or independent review. The concern is not that the authors acted in bad faith; rather, the methodology as described does not rule out the possibility that valid hard questions were removed, causing score inflation and model-dependent bias. The proposed concrete test—independent expert review of a sample of excluded questions—would directly settle whether the filtering was justified. If the test shows the excluded questions are valid, the benchmark's reported accuracies and rankings could change materially, and the central claim would need to be qualified. Therefore, conditional acceptance is appropriate: the paper should be accepted with the requirement that the authors release the excluded questions and provide evidence of independent validation, or address the concern in a revised version. This matches the reader's conditional verdict and confidence.","tokens_in":23771,"tokens_out":3522,"duration_ms":41520,"concrete_test":"Request from the authors the full list of excluded questions (or confirm they are not in the public repository). Draw a random sample of at least 50 excluded questions. Have two or more independent pedagogy experts, who are not authors and are blind to the fact that these questions were excluded, review them using the same criteria described in Section 3.2.3 (relevance, clarity, context-agnostic, and correctness). If a substantial fraction (e.g., >20%) are judged to be valid, unambiguous MCQs, then the filtering step removed valid questions and the reported accuracies are likely inflated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim is that The Pedagogy Benchmark provides a valid, open, and replicable measure of LLM pedagogical knowledge. Validity depends on the dataset construction, particularly the post-hoc filtering in Section 3.2.4. There, approximately 10% of MCQs were excluded because 31 of 36 LLMs answered them consistently incorrectly, with some never answered correctly. The authors state that manual inspection revealed flaws such as grammatical errors, multiple correct answers, or ambiguity. However, the premise that 'consistently incorrect by most LLMs' implies flawed questions is not self-evident. These questions could be valid but genuinely hard, and their removal changes the benchmark's composition. If they are merely hard, reported accuracies are inflated (because hard questions are removed), and the benchmark is biased toward models that share the failure modes of the 36 filtering LLMs. The paper does not report how many of the excluded questions were judged flawed versus simply hard, nor does it provide independent expert verification. This concern is load-bearing because the final dataset composition drives all reported scores, model rankings, and the value-frontier analysis; a biased dataset would undermine the benchmark's validity as a general measure of pedagogical knowledge.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces The Pedagogy Benchmark, a multiple-choice question dataset for evaluating large language models on pedagogical knowledge, sourced from the Chilean Ministry of Education's ECEP teacher exams. The dataset consists of 920 Cross-Domain Pedagogical Knowledge (CDPK) questions and 223 Special Educational Needs and Disabilities (SEND) questions, processed through OCR, translation, de-duplication, expert curation, and a post-hoc filtering step. The authors evaluate 97 LLMs, reporting accuracy scores with bootstrap confidence intervals, subject-level breakdowns, SEND-specific results, and Pareto frontiers for accuracy versus inference cost and model size. They argue that the benchmark fills a gap in education-focused LLM evaluation and can guide model selection for educational applications.","tokens_in":24103,"tokens_out":8231,"duration_ms":86997,"significance":"If the dataset construction is valid, this is a useful and timely contribution: the benchmark is open, the evaluation pipeline is reproducible, and the authors provide a stability test (20 runs, maximum standard deviation 0.43%), bootstrap confidence intervals, and an answer-position bias check. The broad coverage of 97 models and the cost/accuracy analysis are practically valuable. However, the central validity claim depends critically on the post-hoc filtering in Section 3.2.4. If the excluded questions are merely hard rather than flawed, all reported accuracies are inflated and the model rankings are biased toward models similar to the 36 LLMs used for filtering. The manuscript also contains an internal inconsistency about whether the filtering step actually changed the dataset size. These issues must be resolved before the benchmark can be accepted as a valid measure of LLM pedagogical knowledge.","major_comments":[{"comment":"The post-hoc exclusion of approximately 10% of MCQs based on 31 of 36 LLMs answering them incorrectly is not adequately justified. The premise that questions most LLMs answer incorrectly are flawed is an assumption, not a demonstrated fact. If these questions are valid but hard, removing them inflates reported accuracies and biases the benchmark toward models that share the failure modes of the filtering LLMs. The paper does not report how many excluded questions were judged flawed versus simply hard, provides no examples of excluded questions, and includes no independent expert verification or inter-rater reliability. Please provide expert evidence for each excluded question, release the excluded set with justifications, or re-run the benchmark without exclusion and report both versions.","section":"Section 3.2.4"},{"comment":"The reported dataset counts are internally inconsistent. The text says approximately 10% of MCQs were excluded, yet the final dataset is again reported as 1143 MCQs (920 CDPK + 223 SEND), and CDPK accuracy is computed over 899 of 920 questions, implying that only the 21 few-shot examples were removed. If 10% of the 920 CDPK questions had actually been excluded, the denominator would be roughly 828–830, not 920. The manuscript must clarify whether the filtering step changed the dataset, how many questions were removed, and which counts in the abstract and results are correct.","section":"Section 3.2.4 vs Section 3.3"},{"comment":"The value-frontier analysis uses only input token costs, despite the text noting that reasoning models charge for hidden thinking traces as output tokens. Since the frontier includes Gemini 2.5 Flash/Pro and Deepseek R1, excluding output costs may overstate the value of reasoning models relative to non-reasoning models. Please report total (input+output) costs for at least a subset of models, or provide a sensitivity analysis showing that the frontier is robust to this choice.","section":"Section 4.1.2"}],"minor_comments":[{"comment":"The use of GPT-4o-mini for OCR and translation while GPT-4o Mini is later evaluated is a potential same-family bias; please clarify how many questions were manually checked per PDF and whether the pedagogy expert's review covered every translated question.","section":"Section 3.2.1 and Appendix B"},{"comment":"The claim that models 'in many cases exceed human performance' should be tempered because the human baseline in Section 4.1.1 is only an aggregate estimate and does not correspond to actual performance on the benchmark questions.","section":"Section 5.2"},{"comment":"The correlation with other knowledge benchmarks (MMLU, GPQA) is not tested; acknowledging this is a good step, but a small empirical comparison would strengthen the claim that the benchmark isolates pedagogical knowledge.","section":"Section 5.1"},{"comment":"Please report the number of models excluded due to the 5% badly-formatted-response threshold, as this affects the generalizability of the model coverage statement.","section":"Section 3.3"},{"comment":"Table 2 lists accuracies without confidence intervals; adding CIs or referencing Figure 5 would help readers assess the significance of rank differences.","section":"Figure 5 and Table 2"}],"recommendation":"major_revision","confidential_remarks":"This is a benchmark paper rather than a methodological advance, so its fit depends on whether the journal welcomes benchmark contributions. The main technical risk is the filtering step in Section 3.2.4, and the internal count inconsistency makes it unclear whether filtering actually occurred. If the authors can resolve that inconsistency and provide independent validation of the excluded questions, I would support publication. I also note that the curation and filtering rely on a single pedagogical expert, which is a limitation that should be explicitly acknowledged in the main text."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely useful benchmark paper. The Pedagogy Benchmark is a public MCQ dataset of 920 CDPK and 223 SEND questions drawn from Chilean teacher licensing exams, with official answer keys, expert curation, and a careful evaluation pipeline. The cost-accuracy analysis over 97 models is the most actionable part for EdTech model selection, and the open data, code, and live leaderboard make it reproducible. That puts it ahead of many benchmark papers in this space.\n\nWhat it does well: the methodology is transparent about sourcing, de-duplication, translation, and expert annotation. The evaluation uses fixed few-shot prompts, a 20-run stability test (max SD 0.43%), bootstrap CIs, and an answer-position bias check. That is a more thorough evaluation than most comparable MCQ benchmarks. The authors also correctly acknowledge the limitation that MCQ tests measure pedagogical knowledge, not teaching practice, and they do not overclaim beyond that.\n\nThe main soft spot is the post-hoc filtering in Section 3.2.4. About 10% of questions were excluded because 31 of 36 LLMs answered them consistently incorrectly, with manual inspection revealing flaws like ambiguity or multiple correct answers. The concern is that some of those questions may have been merely hard, not flawed. If so, reported accuracies are inflated and the benchmark is biased toward the failure modes of the filtering models. The authors do mention manual inspection, but they do not report how many excluded questions fell into each flaw category, nor do they provide independent expert verification of the excluded set. That is a real validity threat, though not disqualifying. It needs to be documented much more thoroughly, with a breakdown of excluded questions and ideally an external expert review of a sample.\n\nTwo smaller concerns: OCR and translation were done with GPT-4o-mini, a model from a family that is later evaluated; the authors check a random sample but translation artifacts are still possible. And the estimated human baseline of ~50% is built from aggregate exam scores rather than question-level human responses, so it is a rough anchor rather than a validation. The paper also does not compare against general-knowledge benchmarks like MMLU, which the authors acknowledge.\n\nBottom line: this paper deserves a serious referee. I would send it to review with a request that Section 3.2.4 be expanded and that the excluded questions be reported with their flaw categories. The benchmark fills a genuine gap, and the evaluation pipeline is solid enough that the paper is worth engaging with.","headline":"A solid, reproducible pedagogy MCQ benchmark with a real gap to fill; the post-hoc filtering step needs more transparency but is not disqualifying.","tokens_in":24539,"tokens_out":2430,"would_cite":true,"duration_ms":27906,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new benchmark measures whether AI knows how to teach, not just what it knows.","keywords":["pedagogy benchmark","pedagogical knowledge","large language models","LLM evaluation","multiple-choice questions","teacher certification exams","special education needs and disability","cost-accuracy frontier"],"falsifier":"Take the roughly 120 excluded questions and give them to a panel of experienced teachers who have not seen the models' answers: if teachers answer a substantial share correctly, or independent experts judge the items valid, the exclusion was unjustified and the reported scores are too high. A second check is to compute the correlation between CDPK scores and a generic content-knowledge benchmark such as MMLU across the same 97 models; a near-perfect correlation would mean the test measures general multiple-choice ability rather than pedagogy, a possibility the paper leaves open.","tokens_in":23582,"feed_emoji":"🎓","tokens_out":11309,"duration_ms":108277,"temperature":0.7,"pith_summary":"Existing LLM benchmarks such as MMLU test what a model knows about a subject, but none test whether it understands how to teach. This paper argues that pedagogical knowledge — the methods and strategies of teaching — is a distinct, measurable capability, and introduces The Pedagogy Benchmark: 1,143 multiple-choice questions drawn from Chilean teacher certification exams, split into a 920-question cross-domain pedagogy set and a 223-question special-education set. Tested on 97 models, accuracy ranges from 28% for the smallest model to 89% for the current leader, and the authors pair each score with inference cost to show what each dollar of compute buys. If the benchmark is valid, it gives educators and developers a cheap, openly reproducible way to choose models for teaching tools, particularly small on-device models for low-resource classrooms.","feed_headline":"New benchmark measures whether AI knows how to teach","feed_subtitle":"Built from real teacher exams, it ranks 97 AI models and tracks which give the most teaching knowledge per dollar.","key_machinery":"The load-bearing object is the dataset itself: 1,143 multiple-choice questions from teacher professional-development exams, converted from Spanish to English, de-duplicated with fuzzy string matching, trimmed by a pedagogy expert for clarity, relevance, and country-specific content, and then filtered once more by dropping roughly 10% of questions that 31 of 36 LLMs answered incorrectly. The evaluation protocol pairs the dataset with a single fixed three-shot prompt and a lenient answer parser so results are comparable across models, and a four-position answer-swap experiment checks for token- and position-based answer bias. The cost analysis rests on the Pareto value frontier, the convex upper hull of models plotted by inference price against accuracy, which identifies the best-performing model at each price point.","core_discovery":"The paper's central claim is that pedagogical knowledge can be isolated from content knowledge and measured with genuine professional teacher-exam questions, and that current LLMs differ dramatically on this measure. The authors extracted, translated, de-duplicated, and expert-annotated multiple-choice questions from the ECEP teacher evaluation exams administered in Chile from 2017 to 2023, keeping 920 questions for the Cross-Domain Pedagogical Knowledge (CDPK) benchmark and 223 for the Special Educational Needs and Disability (SEND) benchmark, and scored every model with one fixed three-shot prompt. Across 97 models, CDPK accuracy spans 28% to 89% and SEND accuracy 29% to 86%, with the highest scores going to reasoning-focused models and the estimated mean score of human trainee teachers around 50%. The paper further claims that plotting accuracy against inference cost and model size exposes a rapidly improving value frontier: at $0.10 per million input tokens, benchmark accuracy rose from 50% in April 2024 to 82% in June 2025.","pith_inferences":["The paper does not report how CDPK correlates with general knowledge benchmarks such as MMLU or GPQA, and it flags this itself; a direct correlation study across the same 97 models would settle whether pedagogical knowledge is being measured distinctly or just general multiple-choice ability.","The benchmark measures pedagogical knowledge, not classroom practice, and the paper cites teacher-domain evidence linking the two but does not test that link for models; a scenario-based evaluation comparing high- and low-scoring models on generated lesson plans would test the connection.","Because all items come from one country's exams in translation, the benchmark probably reflects that system's pedagogical values; parallel benchmarks built from other countries' teacher exams would show whether the measured 'pedagogical knowledge' is universal or culturally specific.","Re-adding the excluded questions and recomputing the leaderboard is a cheap stability check that would quantify how much the rankings depend on the 10% of items dropped after the model-based filter."],"forward_implications":["Education developers can choose models by budget: open-weight models around 8B parameters score roughly 73% on CDPK, and the paper notes that the April 2024 leader's accuracy is nearly matched at over 400 times lower inference cost.","Pedagogical knowledge is improving rapidly at every price level; at $0.10 per million input tokens, accuracy rose from 50% in April 2024 to 70% in November 2024 to 82% in June 2025.","Reasoning-focused models dominate the top of the leaderboard, indicating that inference-time reasoning improves performance even on a knowledge-retrieval task.","SEND and CDPK scores are highly correlated ($r=0.94$), yet SEND is harder for the best models, so special-education pedagogy is a distinct capability worth evaluating separately.","Small open-weight models that can run on consumer hardware already exceed the estimated 50% human-trainee average, making offline teaching aids feasible in low-connectivity settings."],"supporting_citations":[{"why":"The content-knowledge benchmark this paper builds on and contrasts with; supplies the exam-sourced multiple-choice evaluation methodology.","marker":"[10]"},{"why":"Defines the taxonomy of LLM abilities that places knowledge utilisation, the category this benchmark tests, among other capabilities.","marker":"[28]"},{"why":"Expert-rated assessment of LLM tutor pedagogical practice that the paper identifies as too costly to scale, motivating a replicable MCQ alternative.","marker":"[11]"},{"why":"Documents the answer-position selection bias in LLMs that the paper tests for with its four-position answer-swap experiment.","marker":"[29]"},{"why":"Teacher-domain evidence that pedagogical knowledge aligns with teaching quality, the premise for treating the benchmark as relevant to practice.","marker":"[21]"},{"why":"Establishes the few-shot prompting paradigm whose practices the benchmark's fixed prompt follows.","marker":"[4]"}],"fun_headline_variants":["New benchmark: AI's teaching knowledge varies from 28% to 89%","Teacher-exam benchmark ranks 97 AIs on pedagogy","AI teaching know-how measured: 97 models, wide range","Benchmark tests LLMs on real teacher-exam questions","Do AIs know how to teach? New benchmark says: it depends"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the roughly 10% of questions removed because most of the 36 filtering models answered them incorrectly were genuinely flawed; if those questions were merely hard but valid, the reported accuracies are inflated and the final dataset is biased toward models that resemble the filters.","fun_headline_variants_meta":{"raw":{"variants":["New benchmark: AI's teaching knowledge varies from 28% to 89%","Teacher-exam benchmark ranks 97 AIs on pedagogy","AI teaching know-how measured: 97 models, wide range","Benchmark tests LLMs on real teacher-exam questions","Do AIs know how to teach? New benchmark says: it depends"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000406,"raw_usage":{"total_tokens":2168,"prompt_tokens":1060,"completion_tokens":1108,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":676,"completion_tokens_details":{"reasoning_tokens":1018}},"tokens_in":676,"tokens_out":1108,"duration_ms":13163,"temperature":1.0,"reasoning_tokens":1018,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:14:44.534742+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the roughly 120 excluded questions and give them to a panel of experienced teachers who have not seen the models' answers: if teachers answer a substantial share correctly, or independent experts judge the items valid, the exclusion was unjustified and the reported scores are too high. A second check is to compute the correlation between CDPK scores and a generic content-knowledge benchmark such as MMLU across the same 97 models; a near-perfect correlation would mean the test measures general multiple-choice ability rather than pedagogy, a possibility the paper leaves open.","supporting_citations":[{"cited_title":"Zheng, H","cited_arxiv_id":null,"evidence_quote":"Documents the answer-position selection bias in LLMs that the paper tests for with its four-position answer-swap experiment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Teacher-domain evidence that pedagogical knowledge aligns with teaching quality, the premise for treating the benchmark as relevant to practice."},{"cited_title":"Sastry, A","cited_arxiv_id":null,"evidence_quote":"Establishes the few-shot prompting paradigm whose practices the benchmark's fixed prompt follows."}],"review_version":1}