Pith. sign in

REVIEW 5 cited by

Rethinking the Foundations for Continual Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.08161 v3 pith:4RWJWEIZ submitted 2025-04-10 cs.LG cs.AI

classification cs.LGcs.AI
keywords learningreinforcementfoundationscontinualtraditionalexpectedformalismview
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the traditional view of reinforcement learning, the agent's goal is to find an optimal policy that maximizes its expected sum of rewards. Once the agent finds this policy, the learning ends. This view contrasts with \emph{continual reinforcement learning}, where learning does not end, and agents are expected to continually learn and adapt indefinitely. Despite the clear distinction between these two paradigms of learning, much of the progress in continual reinforcement learning has been shaped by foundations rooted in the traditional view of reinforcement learning. In this paper, we first examine whether the foundations of traditional reinforcement learning are suitable for the continual reinforcement learning paradigm. We identify four key pillars of the traditional reinforcement learning foundations that are antithetical to the goals of continual learning: the Markov decision process formalism, the focus on atemporal artifacts, the expected sum of rewards as an evaluation metric, and episodic benchmark environments that embrace the other three foundations. We then propose a new formalism that sheds the first and the third foundations and replaces them with the history process as a mathematical formalism and a new definition of deviation regret, adapted for continual learning, as an evaluation metric. Finally, we discuss possible approaches to shed the other two foundations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Language Movement Primitives: Grounding Language Models in Robot Motion

    cs.RO 2026-02 conditional novelty 6.0 of 10

    LMP lets a vision-language model set DMP weights and goals from a language command, achieving zero-demonstration tabletop manipulation with reported 80% success on 20 tasks.

  2. Revisiting Adam for Streaming Reinforcement Learning

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    C51 matches StreamQ in streaming RL on 55 Atari games while a new Adaptive Q(λ) algorithm based on bounded derivatives and variance-adjusted updates reaches nearly double the human baseline.

  3. Curriculum-Adapted Robust Reinforcement Learning for UAV Deconfliction in Adversarial Environments

    cs.LG 2025-06 reject novelty 5.0 of 10

    A curriculum that aligns temporal-difference error distributions across increasing adversarial perturbations is claimed to make UAV policies robust to unseen GNSS spoofing attacks, with a generalization certificate.

  4. Darwin Mobile Agent: A Roadmap for Self-Evolution

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    Introduces an open-source mobile GUI agent training framework and a roadmap for autonomous self-evolution via removal of human priors in three pillars.

  5. Agent-centric learning: from external reward maximization to internal knowledge curation

    cs.LG 2025-07 conditional novelty 4.0 of 10

    The paper introduces representational empowerment, a mutual information objective that rewards agents for having internally diverse and controllable representations, as an alternative to external reward maximization.

Pith tools