Pith. sign in

Manifold-aware exploration for reinforcement learning in video generation

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it

fields

cs.CV 4 cs.AI 1

years

2026 5

representative citing papers

CreFlow: Corrective Reflow for Sparse-Reward Embodied Video Diffusion RL

cs.CV · 2026-05-14 · conditional · novelty 7.0

CreFlow combines LTL compositional rewards with credit-aware NFT and corrective reflow losses in online RL to improve embodied video diffusion models, raising downstream task success by 23.8 percentage points on eight bimanual manipulation tasks.

Reward as An Agent for Embodied World Models

cs.AI · 2026-06-18 · conditional · novelty 6.0

Agentic VLM rewards plus dynamic-aware GRPO rollouts let embodied world models explore more diversely while cutting reward hacking and lifting domain scores on Cosmos-Predict2.5 and Kairos3.0-Robot.

citing papers explorer

Showing 5 of 5 citing papers.