Self-improving reactive agents based on reinforcement learning, planning and teaching

Long-Ji Lin · 1992 · DOI 10.1007/bf00992699

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

open at publisher browse 4 citing papers

citation-role summary

background 1

citation-polarity summary

background 1

representative citing papers

Rollout-Level Advantage-Prioritized Experience Replay for GRPO

cs.LG · 2026-06-03 · conditional · novelty 6.0

Rollout-level advantage-prioritized experience replay for GRPO recycles high-advantage individual rollouts with age eviction and fresh-anchored batches to outperform standard GRPO on math benchmarks, with gains increasing with model size.

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning

cs.LG · 2026-06-03 · unverdicted · novelty 5.0

Eligibility traces in deep RL create a peak bias by amplifying distal TD errors into gradient shocks that fixed-step SGD cannot normalize, leading to overestimation of peak-reward trajectories and a mechanistic account of the peak-end rule.

Artifacts as Memory Beyond the Agent Boundary

cs.AI · 2026-04-09 · unverdicted · novelty 5.0

Artifacts in the environment can reduce the memory an RL agent needs to represent its history, as shown by a mathematical proof and experiments with spatial paths.

CLaaS: Continual learning as a service for sample efficient online learning

cs.LG · 2026-06-04 · unverdicted · novelty 4.0

CLaaS enables sample-efficient online continual learning for agents via replay-buffered parametric updates, outperforming in-context learning in forward transfer and retention on an adversarial task.

citing papers explorer

Showing 3 of 3 citing papers after filters.

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning cs.LG · 2026-06-03 · unverdicted · none · ref 43
Eligibility traces in deep RL create a peak bias by amplifying distal TD errors into gradient shocks that fixed-step SGD cannot normalize, leading to overestimation of peak-reward trajectories and a mechanistic account of the peak-end rule.
Artifacts as Memory Beyond the Agent Boundary cs.AI · 2026-04-09 · unverdicted · none · ref 37
Artifacts in the environment can reduce the memory an RL agent needs to represent its history, as shown by a mathematical proof and experiments with spatial paths.
CLaaS: Continual learning as a service for sample efficient online learning cs.LG · 2026-06-04 · unverdicted · none · ref 14
CLaaS enables sample-efficient online continual learning for agents via replay-buffered parametric updates, outperforming in-context learning in forward transfer and retention on an adversarial task.

Self-improving reactive agents based on reinforcement learning, planning and teaching

citation-role summary

citation-polarity summary

fields

years

verdicts

roles

polarities

representative citing papers

citing papers explorer