Pith. sign in

REVIEW 3 cited by

Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short Delays

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.03141 v2 pith:VH637U7T submitted 2024-02-05 cs.LG cs.AIcs.SYeess.SY

classification cs.LGcs.AIcs.SYeess.SY
keywords delaysad-rllearningperformancereinforcementshortstochasticauxiliary
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement learning (RL) is challenging in the common case of delays between events and their sensory perceptions. State-of-the-art (SOTA) state augmentation techniques either suffer from state space explosion or performance degeneration in stochastic environments. To address these challenges, we present a novel Auxiliary-Delayed Reinforcement Learning (AD-RL) method that leverages auxiliary tasks involving short delays to accelerate RL with long delays, without compromising performance in stochastic environments. Specifically, AD-RL learns a value function for short delays and uses bootstrapping and policy improvement techniques to adjust it for long delays. We theoretically show that this can greatly reduce the sample complexity. On deterministic and stochastic benchmarks, our method significantly outperforms the SOTAs in both sample efficiency and policy performance. Code is available at https://github.com/QingyuanWuNothing/AD-RL.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Structural Equivalence and Learning Dynamics in Delayed MARL

    cs.LG 2026-05 accept novelty 8.0 of 10

    Observation and action delays are formally equivalent in cooperative Dec-POMDPs, yielding identical optimal solutions and enabling zero-shot transfer, though learning dynamics differ due to credit assignment and opera...

  2. Diffusion-Guided Uncertainty-Aware Delayed Policy Optimization

    cs.AI 2026-07 conditional novelty 6.0 of 10

    DUPO models the multi-modal posterior of the true state given delayed messages with a diffusion model and uncertainty-weights SAC policy updates, outperforming point-estimate and augmentation baselines under random Mu...

  3. Model-Based Reinforcement Learning under Random Observation Delays

    cs.LG 2025-09 unverdicted novelty 6.0 of 10

    A delay-aware model-based RL framework with sequential belief filtering handles random out-of-sequence observations in POMDPs and outperforms MDP baselines while showing robustness to delay shifts.

Pith tools