pith. sign in

LaSeR: Reinforcement Learning with Last-Token Self-Rewarding

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

fields

cs.AI 1 cs.CL 1

years

2026 2

verdicts

UNVERDICTED 2

representative citing papers

Sparse Reward Subsystem in Large Language Models

cs.CL · 2026-02-01 · unverdicted · novelty 6.0

LLM hidden states contain a sparse reward subsystem consisting of value neurons that predict state value and dopamine neurons that encode step-level temporal difference errors.

citing papers explorer

Showing 2 of 2 citing papers.