Pith. sign in

Representation-based exploration for language models: From test-time to post-training.arXiv preprint arXiv:2510.11686

6 Pith papers cite this work. Polarity classification is still indexing.

6 Pith papers citing it

fields

cs.LG 6

years

2026 6

verdicts

UNVERDICTED 6

representative citing papers

On Advantage Estimates for Max@K Policy Gradients

cs.LG · 2026-06-04 · unverdicted · novelty 6.0

Proposes MaxPO using a Leave-Two-Out baseline for centered unbiased advantages in max@K policy gradients, with a unified derivation of finite-batch estimators.

The Role of Generator Access in Autoregressive Post-Training

cs.LG · 2026-04-06 · unverdicted · novelty 5.0

Limited generator access in autoregressive post-training confines learners to root-start rollouts whose value is bounded by on-policy prefix probabilities, while weak prefix control unlocks richer observations and produces an exponential gap in KL-regularized outcome-reward training.

citing papers explorer

Showing 6 of 6 citing papers.