Pith. sign in

REVIEW 1 cited by

L-SA: Learning Under-Explored Targets in Multi-Target Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.13741 v1 pith:UST4TBA7 submitted 2023-05-23 cs.LG cs.AI

classification cs.LGcs.AI
keywords targetslearningtasksactiveadaptivel-saqueryingsampling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Tasks that involve interaction with various targets are called multi-target tasks. When applying general reinforcement learning approaches for such tasks, certain targets that are difficult to access or interact with may be neglected throughout the course of training - a predicament we call Under-explored Target Problem (UTP). To address this problem, we propose L-SA (Learning by adaptive Sampling and Active querying) framework that includes adaptive sampling and active querying. In the L-SA framework, adaptive sampling dynamically samples targets with the highest increase of success rates at a high proportion, resulting in curricular learning from easy to hard targets. Active querying prompts the agent to interact more frequently with under-explored targets that need more experience or exploration. Our experimental results on visual navigation tasks show that the L-SA framework improves sample efficiency as well as success rates on various multi-target tasks with UTP. Also, it is experimentally demonstrated that the cyclic relationship between adaptive sampling and active querying effectively improves the sample richness of under-explored targets and alleviates UTP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Sparse to Dense: Toddler-inspired Reward Transition in Goal-Oriented Reinforcement Learning

    cs.LG 2025-01 conditional novelty 4.0 of 10

    Sparse-to-dense reward transitions, inspired by toddler learning, improve goal-oriented RL performance and generalization over fixed dense or sparse rewards.

Pith tools