Pith. sign in

The learned critic is a poor training reward in every quadrant; its only robust use is offline diagnosis at the intent frontier

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

RLVP: Penalize the Path, Reward the Outcome

cs.LG · 2026-07-08 · conditional · novelty 6.0

Pairing outcome rewards with verifiable per-action path penalties reduces constraint violations nearly sixfold at equal task success, while a progress potential accelerates learning only where partial progress is reachable.

citing papers explorer

Showing 1 of 1 citing paper.

  • RLVP: Penalize the Path, Reward the Outcome cs.LG · 2026-07-08 · conditional · none · ref 39

    Pairing outcome rewards with verifiable per-action path penalties reduces constraint violations nearly sixfold at equal task success, while a progress potential accelerates learning only where partial progress is reachable.