A lightweight RL policy over a 25-dimensional latent student state and four high-level tutor actions improves simulated tutoring success rates over prompt engineering, but fails to generalize to new math problems.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Efficient RL for optimizing conversation level outcomes with an LLM-based tutor
A lightweight RL policy over a 25-dimensional latent student state and four high-level tutor actions improves simulated tutoring success rates over prompt engineering, but fails to generalize to new math problems.