QuickLAP combines physical corrections with LLM-parsed natural language in a closed-form Bayesian update, reducing reward-learning error in simulated driving and improving user ratings in a 15-person study.
Mehta and Dylan P
2 Pith papers cite this work, alongside 19 external citations. Polarity classification is still indexing.
verdicts
CONDITIONAL 2representative citing papers
A hierarchical teaching algorithm selects complementary environments and feedback modalities to learn reward functions that generalize across unseen MDPs, proving that single-environment teaching leaves structural reward ambiguity.
citing papers explorer
-
QuickLAP: Quick Language-Action Preference Learning for Semi-Autonomous Agents
QuickLAP combines physical corrections with LLM-parsed natural language in a closed-form Bayesian update, reducing reward-learning error in simulated driving and improving user ratings in a 15-person study.
-
Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning
A hierarchical teaching algorithm selects complementary environments and feedback modalities to learn reward functions that generalize across unseen MDPs, proving that single-environment teaching leaves structural reward ambiguity.