PairGS builds a relation graph from sparse pairwise affinities on 3D Gaussians to achieve SOTA open-vocabulary segmentation with a 50x faster variant than optimization-based methods.
arXiv preprint arXiv:2507.13097 (2025)
9 Pith papers cite this work. Polarity classification is still indexing.
abstract
Grasping is a fundamental robot skill, yet despite significant research advancements, learning-based 6-DOF grasping approaches are still not turnkey and struggle to generalize across different embodiments and in-the-wild settings. We build upon the recent success on modeling the object-centric grasp generation process as an iterative diffusion process. Our proposed framework, GraspGen, consists of a DiffusionTransformer architecture that enhances grasp generation, paired with an efficient discriminator to score and filter sampled grasps. We introduce a novel and performant on-generator training recipe for the discriminator. To scale GraspGen to both objects and grippers, we release a new simulated dataset consisting of over 53 million grasps. We demonstrate that GraspGen outperforms prior methods in simulations with singulated objects across different grippers, achieves state-of-the-art performance on the FetchBench grasping benchmark, and performs well on a real robot with noisy visual observations.
representative citing papers
Real-IKEA supplies 1,079 physically accurate articulated asset configurations from real IKEA parts together with resistance-calibrated simulation parameters that enable RL policies to discover robust hooking and levering behaviors.
VoLoAgent uses a VLM to steer heterogeneous robot capabilities as interruptible tools for long-horizon manipulation and introduces the RoboVoLo benchmark, claiming substantial outperformance over single VLA/VLM or tool-based systems with real-robot validation.
GraspGen-X extends diffusion 6-DOF grasping to cross-embodiment via swept-volume gripper encoding, trained on procedural grippers and 2B grasps, claiming best zero-shot generalization to novel grippers in sim and real tests.
GEM-4D improves video world models for robot manipulation by distilling 4D geometric correspondences into training and adding an inverse dynamics module, achieving SOTA geometric consistency and 81% real-world success.
IGen generates realistic visuomotor training data including actions and temporally coherent visuals from unstructured open-world images via 3D reconstruction and VLM reasoning.
A unified generative pipeline produces cross-simulator, affordance-annotated, task-conditioned 3D worlds that support online robot policy training and real-robot transfer.
Training-free Sim(3) alignment method with hallucination filtering for generative 3D to partial monocular registration, plus new GenPMOAlign benchmark showing outperformance over classical and learning baselines.
Human2Any transfers human video demonstrations to robots by representing tasks as object-object interactions and composing learned priors with robot-side planning.
citing papers explorer
-
Relation-Centric Open-Vocabulary 3D Gaussian Segmentation
PairGS builds a relation graph from sparse pairwise affinities on 3D Gaussians to achieve SOTA open-vocabulary segmentation with a 50x faster variant than optimization-based methods.
-
Real-IKEA: Physical Fidelity is the Prerequisite for Robust Manipulation
Real-IKEA supplies 1,079 physically accurate articulated asset configurations from real IKEA parts together with resistance-calibrated simulation parameters that enable RL policies to discover robust hooking and levering behaviors.
-
VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation
VoLoAgent uses a VLM to steer heterogeneous robot capabilities as interruptible tools for long-horizon manipulation and introduces the RoboVoLo benchmark, claiming substantial outperformance over single VLA/VLM or tool-based systems with real-robot validation.
-
GraspGen-X: Cross-Embodiment 6-DOF Diffusion-based Grasping
GraspGen-X extends diffusion 6-DOF grasping to cross-embodiment via swept-volume gripper encoding, trained on procedural grippers and 2B grasps, claiming best zero-shot generalization to novel grippers in sim and real tests.
-
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
GEM-4D improves video world models for robot manipulation by distilling 4D geometric correspondences into training and adding an inverse dynamics module, achieving SOTA geometric consistency and 81% real-world success.
-
IGen: Scalable Data Generation for Robot Learning from Open-World Images
IGen generates realistic visuomotor training data including actions and temporally coherent visuals from unstructured open-world images via 3D reconstruction and VLM reasoning.
-
EmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AI
A unified generative pipeline produces cross-simulator, affordance-annotated, task-conditioned 3D worlds that support online robot policy training and real-robot transfer.
-
Robust 3D Alignment of Generative Reconstructions via Partial Monocular Observations
Training-free Sim(3) alignment method with hallucination filtering for generative 3D to partial monocular registration, plus new GenPMOAlign benchmark showing outperformance over classical and learning baselines.
-
Human2Any: Human-to-Robot Transfer via Constraint-Aware Compositional Planning
Human2Any transfers human video demonstrations to robots by representing tasks as object-object interactions and composing learned priors with robot-side planning.