Pith. sign in

Enhancing vision- language model training with reinforcement learning in synthetic worlds for real-world success

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

citation-role summary

background 1 method 1

citation-polarity summary

years

2026 3

verdicts

UNVERDICTED 3

representative citing papers

Rank-Then-Act: Reward-Free Control from Frame-Order Progress

cs.LG · 2026-07-02 · unverdicted · novelty 6.0

RTA trains a VLM as a progress ordinal scorer via GRPO on shuffled expert frames and uses Spearman rank correlation with temporal indices as a bounded RL reward, matching or exceeding prior video reward methods on discrete and continuous control benchmarks.

citing papers explorer

Showing 3 of 3 citing papers.