Pith. sign in

REVIEW 2 cited by

Opinion-Guided Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.17287 v2 pith:RI454OYJ submitted 2024-05-27 cs.LG cs.AI

classification cs.LGcs.AI
keywords learningopinionsreinforcementagentsapproachhumanadvisorshigher
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Human guidance is often desired in reinforcement learning to improve the performance of the learning agent. However, human insights are often mere opinions and educated guesses rather than well-formulated arguments. While opinions are subject to uncertainty, e.g., due to partial informedness or ignorance about a problem, they also emerge earlier than hard evidence can be produced. Thus, guiding reinforcement learning agents by way of opinions offers the potential for more performant learning processes, but comes with the challenge of modeling and managing opinions in a formal way. In this article, we present a method to guide reinforcement learning agents through opinions. To this end, we provide an end-to-end method to model and manage advisors' opinions. To assess the utility of the approach, we evaluate it with synthetic (oracle) and human advisors, at different levels of uncertainty, and under multiple advice strategies. Our results indicate that opinions, even if uncertain, improve the performance of reinforcement learning agents, resulting in higher rewards, more efficient exploration, and a better reinforced policy. Although we demonstrate our approach through a two-dimensional topological running example, our approach is applicable to complex problems with higher dimensions as well.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AI Simulation by Digital Twins: Systematic Survey, Reference Framework, and Mapping to a Standardized Architecture

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A systematic survey of digital twin enabled AI simulation produces the DT4AI reference framework and maps it to ISO 23247.

  2. Complex Model Transformations by Reinforcement Learning with Uncertain Human Guidance

    cs.SE 2025-06 conditional novelty 5.0 of 10

    A framework that fuses uncertain human advice, encoded as subjective logic opinions, into the policy of a reinforcement learning agent improves learning of model transformation sequences, but only convincingly at low-...

Pith tools