Pith. sign in

REVIEW 1 cited by

Quasimetric Value Functions with Dense Rewards

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.08724 v1 pith:WKESWXSX submitted 2024-09-13 cs.LG cs.AI

classification cs.LGcs.AI
keywords densegcrlquasimetricrewardrewardscomplexitysamplesetting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

As a generalization of reinforcement learning (RL) to parametrizable goals, goal conditioned RL (GCRL) has a broad range of applications, particularly in challenging tasks in robotics. Recent work has established that the optimal value function of GCRL $Q^\ast(s,a,g)$ has a quasimetric structure, leading to targetted neural architectures that respect such structure. However, the relevant analyses assume a sparse reward setting -- a known aggravating factor to sample complexity. We show that the key property underpinning a quasimetric, viz., the triangle inequality, is preserved under a dense reward setting as well. Contrary to earlier findings where dense rewards were shown to be detrimental to GCRL, we identify the key condition necessary for triangle inequality. Dense reward functions that satisfy this condition can only improve, never worsen, sample complexity. This opens up opportunities to train efficient neural architectures with dense rewards, compounding their benefits to sample complexity. We evaluate this proposal in 12 standard benchmark environments in GCRL featuring challenging continuous control tasks. Our empirical results confirm that training a quasimetric value function in our dense reward setting indeed outperforms training with sparse rewards.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PPAAS: PVT and Pareto Aware Analog Sizing via Goal-conditioned Reinforcement Learning

    eess.SP 2025-07 conditional novelty 6.0 of 10

    PPAAS uses a goal-conditioned Soft Actor-Critic policy with Pareto-front goal sampling and conservative hindsight replay to improve PVT-aware analog circuit sizing efficiency.

Pith tools