Pith. sign in

REVIEW 14 cited by

FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.08261 v2 pith:3FIINL5B submitted 2024-12-11 cs.RO cs.AIcs.LG

FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World Model

classification cs.RO cs.AIcs.LG
keywords modelvideoflipflowlong-horizonplanninggeneral-purposegeneration
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We aim to develop a model-based planning framework for world models that can be scaled with increasing model and data budgets for general-purpose manipulation tasks with only language and vision inputs. To this end, we present FLow-centric generative Planning (FLIP), a model-based planning algorithm on visual space that features three key modules: 1. a multi-modal flow generation model as the general-purpose action proposal module; 2. a flow-conditioned video generation model as the dynamics module; and 3. a vision-language representation learning model as the value module. Given an initial image and language instruction as the goal, FLIP can progressively search for long-horizon flow and video plans that maximize the discounted return to accomplish the task. FLIP is able to synthesize long-horizon plans across objects, robots, and tasks with image flows as the general action representation, and the dense flow information also provides rich guidance for long-horizon video generation. In addition, the synthesized flow and video plans can guide the training of low-level control policies for robot execution. Experiments on diverse benchmarks demonstrate that FLIP can improve both the success rates and quality of long-horizon video plan synthesis and has the interactive world model property, opening up wider applications for future works.Video demos are on our website: https://nus-lins-lab.github.io/flipweb/.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization

    cs.CV 2026-06 unverdicted novelty 7.0

    VLMs act as teachers by deriving differentiable rewards from task rules to adapt VGMs via test-time LoRA optimization, delivering 16.7-point average gains on symbolic and general video reasoning benchmarks.

  2. VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization

    cs.CV 2026-06 unverdicted novelty 7.0

    VLMs formulate differentiable rewards from task-specific rules to enable test-time online LoRA optimization of VGMs, delivering 16.7-point gains on symbolic and general video reasoning benchmarks over VLM-as-solver an...

  3. AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation

    cs.RO 2026-05 unverdicted novelty 7.0

    AwareVLN introduces a structural reasoning module and automatic data engine with progress division to equip VLN agents with self-awareness of agent state and task progress, outperforming prior methods on Habitat datasets.

  4. RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic Manipulation

    cs.RO 2026-05 unverdicted novelty 7.0

    RoboFlow4D is an end-to-end lightweight flow world model that predicts multi-frame 3D flows from visual observations and textual instructions to provide explicit planning for real-time robotic manipulation.

  5. SSI-Policy: Learning Structured Scene Interfaces for Vision-Language Robotic Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0

    SSI-Policy learns a robot-agnostic RGB-only scene interface from video to improve vision-language manipulation policies by 15% on LIBERO with only 10 demos per task.

  6. SSI-Policy: Learning Structured Scene Interfaces for Vision-Language Robotic Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0

    SSI-Policy uses an RGB-only Structured Scene Interface to improve LIBERO benchmark performance by nearly 15% with only 10 demonstrations per task compared to prior methods.

  7. Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation

    cs.CV 2025-12 conditional novelty 6.0

    Reward Forcing combines EMA-Sink tokens and Rewarded Distribution Matching Distillation to deliver state-of-the-art streaming video generation at 23.1 FPS without copying initial frames.

  8. Ctrl-World: A Controllable Generative World Model for Robot Manipulation

    cs.RO 2025-10 unverdicted novelty 6.0

    A controllable world model trained on the DROID dataset generates consistent multi-view robot trajectories for over 20 seconds and improves generalist policy success rates by 44.7% via imagined trajectory fine-tuning.

  9. F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions

    cs.RO 2025-09 unverdicted novelty 6.0

    F1 integrates next-scale visual foresight prediction into a Mixture-of-Transformer VLA architecture to reformulate action generation as foresight-guided inverse dynamics, achieving higher success rates on 136 tasks.

  10. Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations

    cs.RO 2025-07 unverdicted novelty 6.0

    RIGVid shows that filtered AI-generated videos can serve as effective supervision for complex robotic manipulation tasks without any real demonstrations.

  11. PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration

    cs.CV 2026-07 conditional novelty 5.0

    PAVXploreRL post-trains action-conditioned world models with VJEPA-2 latent rewards and perturbed 'OOD' actions, reporting a 5.6% average gain and lowered policy-overestimation bias.

  12. A Survey on Vision-Language-Action Models: An Action Tokenization Perspective

    cs.RO 2025-07 unverdicted novelty 5.0

    The survey frames VLA models as pipelines that generate progressively grounded action tokens and classifies those tokens into eight types to guide future development.

  13. Position: Good Embodied Reward Models Need Bad Behavior Data

    cs.RO 2026-05 unverdicted novelty 4.0

    Embodied reward models systematically over-reward unsafe, suboptimal, and shortcut robot behaviors due to training on successful data only, and modest inclusion of bad behavior data improves alignment with human preferences.

  14. Agentic Reasoning for Large Language Models

    cs.AI 2026-01 unverdicted novelty 4.0

    The survey structures agentic reasoning for LLMs into foundational, self-evolving, and collective multi-agent layers while distinguishing in-context orchestration from post-training optimization and reviewing applicat...