pith. sign in

hub

RL-VLM-F: Reinforcement learn- ing from vision language foundation model feedback

13 Pith papers cite this work. Polarity classification is still indexing.

13 Pith papers citing it

hub tools

citation-role summary

background 1

citation-polarity summary

years

2026 11 2025 2

verdicts

UNVERDICTED 13

roles

background 1

polarities

background 1

clear filters

representative citing papers

Freeform Preference Learning for Robotic Manipulation

cs.RO · 2026-06-30 · unverdicted · novelty 6.0

Freeform Preference Learning trains language-conditioned multi-axis reward models from human pairwise preferences to produce steerable and compositional robot policies that outperform sparse and binary-preference baselines by 38 percentage points.

Reflection-Based Task Adaptation for Self-Improving VLA

cs.RO · 2025-10-14 · unverdicted · novelty 5.0

Reflective Self-Adaptation combines failure-reflective reinforcement learning with success-guided imitation learning to enable faster and more reliable task adaptation for pre-trained Vision-Language-Action models.

citing papers explorer

Showing 13 of 13 citing papers after filters.