Pith. sign in

REVIEW 2 cited by

A Practical Guide to Multi-Objective Reinforcement Learning and Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.09568 v1 pith:U4CMGOZC submitted 2021-03-17 cs.AI cs.LG

classification cs.AIcs.LG
keywords multi-objectivelearningplanningproblemsreinforcementcomplexdecision-makingguide
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Real-world decision-making tasks are generally complex, requiring trade-offs between multiple, often conflicting, objectives. Despite this, the majority of research in reinforcement learning and decision-theoretic planning either assumes only a single objective, or that multiple objectives can be adequately handled via a simple linear combination. Such approaches may oversimplify the underlying problem and hence produce suboptimal results. This paper serves as a guide to the application of multi-objective methods to difficult problems, and is aimed at researchers who are already familiar with single-objective reinforcement learning and planning methods who wish to adopt a multi-objective perspective on their research, as well as practitioners who encounter multi-objective decision problems in practice. It identifies the factors that may influence the nature of the desired solution, and illustrates by example how these influence the design of multi-objective decision-making systems for complex problems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Test-Time Scaling via Error Localization

    cs.LG 2026-07 conditional novelty 6.0 of 10

    TTEL uses feedback-induced token probability drops to localize the first error in a failed reasoning trace and branch a new generation from that prefix, improving pass@k per token on coding and math benchmarks.

  2. CuRLA: Curriculum Learning Based Deep Reinforcement Learning for Autonomous Driving

    cs.RO 2025-01 conditional novelty 4.0 of 10

    A curriculum that progressively adds traffic and a collision penalty, plus a speed-encouraging reward, increases the average speed of a PPO+VAE driving agent in CARLA while keeping lap distance comparable to a baseline.

Pith tools