In simulated five-round teaching dialogues, Llama 3.1 70B Instruct produced the largest pre/post accuracy gains among 14 LLMs, and teaching effectiveness did not track model scale or benchmark reasoning scores.
Duschl and Drew H
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
EducationQ: Evaluating LLMs' Teaching Capabilities Through Multi-Agent Dialogue Framework
In simulated five-round teaching dialogues, Llama 3.1 70B Instruct produced the largest pre/post accuracy gains among 14 LLMs, and teaching effectiveness did not track model scale or benchmark reasoning scores.