Best submitted systems scored 58-72 macro F1 on four three-class pedagogical assessment tracks and 97 macro F1 on nine-class tutor identification, showing automatic evaluation of AI math tutors works but still has room to improve.
In Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CY 1years
2025 1verdicts
ACCEPT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors
Best submitted systems scored 58-72 macro F1 on four three-class pedagogical assessment tracks and 97 macro F1 on nine-class tutor identification, showing automatic evaluation of AI math tutors works but still has room to improve.