Pith. sign in

hub

TOPRe- ward: Token probabilities as hidden zero-shot rewards for robotics

14 Pith papers cite this work. Polarity classification is still indexing.

14 Pith papers citing it

hub tools

citation-role summary

background 3

citation-polarity summary

years

2026 14

roles

background 3

polarities

background 3

representative citing papers

Improving Robotic Generalist Policies via Flow Reversal Steering

cs.RO · 2026-06-11 · unverdicted · novelty 7.0

Flow Reversal Steering steers flow matching generalist policies by reversing suboptimal actions to nearby better modes, enabling improved zero-shot control, quick distillation, and RL bootstrapping in robotic manipulation.

Freeform Preference Learning for Robotic Manipulation

cs.RO · 2026-06-30 · conditional · novelty 6.0

FPL trains a language-conditioned reward model from per-axis human preferences and a reward-conditioned policy, reporting 38-point average success gains over sparse-reward and binary-preference baselines on six manipulation tasks.

LLM-as-a-Verifier: A General-Purpose Verification Framework

cs.AI · 2026-07-06 · conditional · novelty 5.0

Expecting over scoring-token logits yields continuous, scalable verification that improves agent trajectory selection and dense RL rewards across coding, robotics, and medical benchmarks.

World Value Models for Robotic Manipulation

cs.RO · 2026-06-23 · unverdicted · novelty 5.0

World Value Model (WVM) integrates world models with value estimation to achieve SOTA Value-Order Correlation on expert and suboptimal robotic data and improves downstream policy performance.

citing papers explorer

Showing 14 of 14 citing papers.