Pith. sign in

REVIEW 2 cited by

Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.02160 v1 pith:4L5BEVFD submitted 2024-01-04 cs.NE

classification cs.NE
keywords morlpoliciesalgorithmsinformationoptimizationpolicypreferencepreference-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-objective reinforcement learning (MORL) aims to find a set of high-performing and diverse policies that address trade-offs between multiple conflicting objectives. However, in practice, decision makers (DMs) often deploy only one or a limited number of trade-off policies. Providing too many diversified trade-off policies to the DM not only significantly increases their workload but also introduces noise in multi-criterion decision-making. With this in mind, we propose a human-in-the-loop policy optimization framework for preference-based MORL that interactively identifies policies of interest. Our method proactively learns the DM's implicit preference information without requiring any a priori knowledge, which is often unavailable in real-world black-box decision scenarios. The learned preference information is used to progressively guide policy optimization towards policies of interest. We evaluate our approach against three conventional MORL algorithms that do not consider preference information and four state-of-the-art preference-based MORL algorithms on two MORL environments for robot control and smart grid management. Experimental results fully demonstrate the effectiveness of our proposed method in comparison to the other peer algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Humans Coexist, So Must Embodied Artificial Agents

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Coexistence, defined as sustained meaningful and reciprocal interaction among an agent, humans, and environment, is presented as a necessary design goal for embodied AI.

  2. Preference-based Multi-Objective Reinforcement Learning

    cs.LG 2025-07 reject novelty 4.0 of 10

    Pb-MORL learns a multi-objective reward model from preference comparisons and claims to achieve Pareto-optimal policies, outperforming an oracle in energy and highway tasks.

Pith tools