REVIEW 11 cited by
This&That: Language-Gesture Controlled Video Generation for Robot Planning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Clear, interpretable instructions are invaluable when attempting any complex task. Good instructions help to clarify the task and even anticipate the steps needed to solve it. In this work, we propose a robot learning framework for communicating, planning, and executing a wide range of tasks, dubbed This&That. This&That solves general tasks by leveraging video generative models, which, through training on internet-scale data, contain rich physical and semantic context. In this work, we tackle three fundamental challenges in video-based planning: 1) unambiguous task communication with simple human instructions, 2) controllable video generation that respects user intent, and 3) translating visual plans into robot actions. This&That uses language-gesture conditioning to generate video predictions, as a succinct and unambiguous alternative to existing language-only methods, especially in complex and uncertain environments. These video predictions are then fed into a behavior cloning architecture dubbed Diffusion Video to Action (DiVA), which outperforms prior state-of-the-art behavior cloning and video-based planning methods by substantial margins.
Forward citations
Cited by 11 Pith papers
-
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
A dual-stream diffusion VLA that predicts actions and future observations in separate streams with shared attention beats VLA/world-model baselines on simulated and real robot tasks.
-
Ego-centric Predictive Model Conditioned on Hand Trajectories
Ego-PM predicts future hand trajectories and then uses them to condition latent diffusion video generation, jointly outputting actions and future frames in egocentric and robotic scenes.
-
RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation
A hierarchical pipeline decomposes long-horizon robot instructions into keyframes, interpolates between them, and regresses joint states from the generated video, reaching 67.4% success on simulated long-horizon tasks.
-
DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation
A diffusion model fuses per-person masks with pose keypoints to generate identity-preserving, two-person interactive videos from a single reference image, outperforming prior single-person-animation pipelines.
-
Solving New Tasks by Adapting Internet Video Knowledge
Inverse Probabilistic Adaptation, a score-composition variant that keeps the large video model as the base and consults a small in-domain model, achieves 68.3% average success on MetaWorld policy supervision and stays...
-
FlowLoss: Dynamic Flow-Conditioned Loss Strategy for Video Diffusion Models
FlowLoss adds a noise-gated optical-flow matching loss to video diffusion training, yielding earlier motion stability but mixed final results and higher training cost.
-
Generative World Explorer
GenEx generates consistent 360-degree exploration videos from a single egocentric panorama and shows that LLM agents using these imagined views make more accurate decisions.
-
GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning
GenFlowRL converts generated 2D object keypoint flows into a compact delta-flow that shapes dense RL rewards, improving robot manipulation performance on 10 simulation and real-world probe tasks.
-
EnerVerse-AC: Envisioning Embodied Environments with Action Condition
EnerVerse-AC generates realistic multi-view robot videos conditioned on action sequences and shows early evidence it can augment training data and rank policy performance like a real robot.
-
GenEx: Generating an Explorable World
GenEx builds a consistent 360-degree explorable world from a single image and uses imagined videos to improve GPT-based embodied decisions.
-
ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance
ManipDreamer conditions a robot-manipulation video diffusion model on action-tree instruction embeddings and multi-modal visual guidance, reporting modest gains over RoboDreamer that are undercut by evaluation inconsi...
Discussion (0). Continue with ORCID to comment.