Pith. sign in

REVIEW 1 cited by

DEER: A Delay-Resilient Framework for Reinforcement Learning with Variable Delays

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.03102 v1 pith:CP3WY3HL submitted 2024-06-05 cs.LG cs.AI

classification cs.LGcs.AI
keywords deeralgorithmsdelaydelaysstateschallengesdelay-resilientdelayed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Classic reinforcement learning (RL) frequently confronts challenges in tasks involving delays, which cause a mismatch between received observations and subsequent actions, thereby deviating from the Markov assumption. Existing methods usually tackle this issue with end-to-end solutions using state augmentation. However, these black-box approaches often involve incomprehensible processes and redundant information in the information states, causing instability and potentially undermining the overall performance. To alleviate the delay challenges in RL, we propose $\textbf{DEER (Delay-resilient Encoder-Enhanced RL)}$, a framework designed to effectively enhance the interpretability and address the random delay issues. DEER employs a pretrained encoder to map delayed states, along with their variable-length past action sequences resulting from different delays, into hidden states, which is trained on delay-free environment datasets. In a variety of delayed scenarios, the trained encoder can seamlessly integrate with standard RL algorithms without requiring additional modifications and enhance the delay-solving capability by simply adapting the input dimension of the original algorithms. We evaluate DEER through extensive experiments on Gym and Mujoco environments. The results confirm that DEER is superior to state-of-the-art RL algorithms in both constant and random delay settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diffusion-Guided Uncertainty-Aware Delayed Policy Optimization

    cs.AI 2026-07 conditional novelty 6.0 of 10

    DUPO models the multi-modal posterior of the true state given delayed messages with a diffusion model and uncertainty-weights SAC policy updates, outperforming point-estimate and augmentation baselines under random Mu...

Pith tools