DP-RL augments RL policy gradients with auxiliary losses from dynamical priors to enforce temporally structured decision trajectories that depend on task-specific evidence accumulation and hysteresis.
Probabilistic decision making by slow reverberation in cortical circuits
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Dynamical Priors as a Training Objective in Reinforcement Learning
DP-RL augments RL policy gradients with auxiliary losses from dynamical priors to enforce temporally structured decision trajectories that depend on task-specific evidence accumulation and hysteresis.