A framework that fuses uncertain human advice, encoded as subjective logic opinions, into the policy of a reinforcement learning agent improves learning of model transformation sequences, but only convincingly at low-to-moderate advice uncertainty.
Opinion-Guided Reinforcement Learning
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Human guidance is often desired in reinforcement learning to improve the performance of the learning agent. However, human insights are often mere opinions and educated guesses rather than well-formulated arguments. While opinions are subject to uncertainty, e.g., due to partial informedness or ignorance about a problem, they also emerge earlier than hard evidence can be produced. Thus, guiding reinforcement learning agents by way of opinions offers the potential for more performant learning processes, but comes with the challenge of modeling and managing opinions in a formal way. In this article, we present a method to guide reinforcement learning agents through opinions. To this end, we provide an end-to-end method to model and manage advisors' opinions. To assess the utility of the approach, we evaluate it with synthetic (oracle) and human advisors, at different levels of uncertainty, and under multiple advice strategies. Our results indicate that opinions, even if uncertain, improve the performance of reinforcement learning agents, resulting in higher rewards, more efficient exploration, and a better reinforced policy. Although we demonstrate our approach through a two-dimensional topological running example, our approach is applicable to complex problems with higher dimensions as well.
citation-role summary
citation-polarity summary
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1roles
other 1polarities
unclear 1representative citing papers
citing papers explorer
-
Complex Model Transformations by Reinforcement Learning with Uncertain Human Guidance
A framework that fuses uncertain human advice, encoded as subjective logic opinions, into the policy of a reinforcement learning agent improves learning of model transformation sequences, but only convincingly at low-to-moderate advice uncertainty.