pith. sign in

Self-improving reactive agents based on reinforcement learning, planning and teaching

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 3 cs.AI 1

years

2026 4

roles

background 1

polarities

background 1

clear filters

representative citing papers

Rollout-Level Advantage-Prioritized Experience Replay for GRPO

cs.LG · 2026-06-03 · conditional · novelty 6.0

Rollout-level advantage-prioritized experience replay for GRPO recycles high-advantage individual rollouts with age eviction and fresh-anchored batches to outperform standard GRPO on math benchmarks, with gains increasing with model size.

Artifacts as Memory Beyond the Agent Boundary

cs.AI · 2026-04-09 · unverdicted · novelty 5.0

Artifacts in the environment can reduce the memory an RL agent needs to represent its history, as shown by a mathematical proof and experiments with spatial paths.

citing papers explorer

Showing 3 of 3 citing papers after filters.