Pith. sign in

REVIEW

Recurrent Reinforcement Learning with Memoroids

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.09900 v3 pith:OJWUHIK2 submitted 2024-02-15 cs.LG cs.AI

classification cs.LGcs.AI
keywords recurrentmodelslearningmemoroidsreinforcementbatchingmarkovmemory
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Memory models such as Recurrent Neural Networks (RNNs) and Transformers address Partially Observable Markov Decision Processes (POMDPs) by mapping trajectories to latent Markov states. Neither model scales particularly well to long sequences, especially compared to an emerging class of memory models called Linear Recurrent Models. We discover that the recurrent update of these models resembles a monoid, leading us to reformulate existing models using a novel monoid-based framework that we call memoroids. We revisit the traditional approach to batching in recurrent reinforcement learning, highlighting theoretical and empirical deficiencies. We leverage memoroids to propose a batching method that improves sample efficiency, increases the return, and simplifies the implementation of recurrent loss functions in reinforcement learning.

Discussion (0). Continue with ORCID to comment.

Pith tools