Instant-Fold enables execution of multiple deformable object manipulation modes from a single demonstration via a flow-matching transformer policy that transfers zero-shot from simulation to real robots.
arXiv preprint arXiv:2508.11002 (2025) 3
6 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 6roles
background 1polarities
background 1representative citing papers
ChronoFlow-Policy improves visuomotor policy learning by co-training action generation with prediction of past-current-future gripper-object keypoint trajectories.
3D HAMSTER adds depth encoding and reconstruction to VLMs to produce 3D waypoint sequences that feed directly into pointcloud policies, claiming better generalization than 2D baselines under shifts.
MV-Actor proposes a multi-view framework using semantic interaction, semantic-spatial token interaction, and guided depth repair to reach 87.8% success on the PerAct2 bimanual benchmark and outperform RGB/RGB-D baselines in real-world tests.
PhysMani couples a physics-principled 3D Gaussian world model with a future-aware policy to achieve higher success rates on dynamic manipulation tasks in simulation and real robots.
A transformer 3D encoder plus diffusion decoder architecture, with 3D-specific augmentations, outperforms prior 3D policy methods on manipulation benchmarks by improving training stability.
citing papers explorer
-
Instant-Fold: In-Context Imitation Learning for Deformable Object Manipulation
Instant-Fold enables execution of multiple deformable object manipulation modes from a single demonstration via a flow-matching transformer policy that transfers zero-shot from simulation to real robots.
-
ChronoFlow-Policy: Unifying Past-Current-Future Interaction Flow in Visuomotor Policy Learning
ChronoFlow-Policy improves visuomotor policy learning by co-training action generation with prediction of past-current-future gripper-object keypoint trajectories.
-
3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance
3D HAMSTER adds depth encoding and reconstruction to VLMs to produce 3D waypoint sequences that feed directly into pointcloud policies, claiming better generalization than 2D baselines under shifts.
-
MV-Actor: Aligning Multi-View Semantics and Spatial Awareness for Bimanual Manipulation
MV-Actor proposes a multi-view framework using semantic interaction, semantic-spatial token interaction, and guided depth repair to reach 87.8% success on the PerAct2 bimanual benchmark and outperform RGB/RGB-D baselines in real-world tests.
-
PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
PhysMani couples a physics-principled 3D Gaussian world model with a future-aware policy to achieve higher success rates on dynamic manipulation tasks in simulation and real robots.
-
R3D: Revisiting 3D Policy Learning
A transformer 3D encoder plus diffusion decoder architecture, with 3D-specific augmentations, outperforms prior 3D policy methods on manipulation benchmarks by improving training stability.