A PPO policy for deciding topic order and duration on a prerequisite knowledge graph, paired with an LLM for Socratic dialogue, improves student mastery rates and reduces turns compared to baselines and scaled models across held-out topics.
InProceedings of the 2024 Joint International Con- ference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 5414–5424, Torino, Italia
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CL 2representative citing papers
DIAL uses iterative adversarial learning to restore lexical diversity in user simulators and achieve strong correlation with real failure rates while keeping low divergence in failure mode distributions for mental health dialogues.
citing papers explorer
-
Hey Chat, Can You Teach Me? Structuring Socratic Dialogue for Human Learning in the Wild
A PPO policy for deciding topic order and duration on a prerequisite knowledge graph, paired with an LLM for Socratic dialogue, improves student mastery rates and reduces turns compared to baselines and scaled models across held-out topics.
-
DIAL: Direct Iterative Adversarial Learning for Realistic Multi-Turn Dialogue Simulation
DIAL uses iterative adversarial learning to restore lexical diversity in user simulators and achieve strong correlation with real failure rates while keeping low divergence in failure mode distributions for mental health dialogues.