Pith. sign in

Leveraging Pretrained Diffusion Models for Zero-Shot Part Assembly

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

3D part assembly aims to understand part relationships and predict their 6-DoF poses to construct realistic 3D shapes, addressing the growing demand for autonomous assembly, which is crucial for robots. Existing methods mainly estimate the transformation of each part by training neural networks under supervision, which requires a substantial quantity of manually labeled data. However, the high cost of data collection and the immense variability of real-world shapes and parts make traditional methods impractical for large-scale applications. In this paper, we propose first a zero-shot part assembly method that utilizes pre-trained point cloud diffusion models as discriminators in the assembly process, guiding the manipulation of parts to form realistic shapes. Specifically, we theoretically demonstrate that utilizing a diffusion model for zero-shot part assembly can be transformed into an Iterative Closest Point (ICP) process. Then, we propose a novel pushing-away strategy to address the overlap parts, thereby further enhancing the robustness of the method. To verify our work, we conduct extensive experiments and quantitative comparisons to several strong baseline methods, demonstrating the effectiveness of the proposed approach, which even surpasses the supervised learning method. The code has been released on https://github.com/Ruiyuan-Zhang/Zero-Shot-Assembly.

fields

eess.AS 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis

eess.AS · 2025-05-20 · conditional · novelty 6.0

TCSinger 2 generates zero-shot singing voices in nine languages with style transfer from audio prompts and multi-level style control from natural language prompts, using blurred boundary encoders, contrastive prompt alignment, and a flow-based transformer.

citing papers explorer

Showing 1 of 1 citing paper.

  • TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis eess.AS · 2025-05-20 · conditional · none · ref 38 · internal anchor

    TCSinger 2 generates zero-shot singing voices in nine languages with style transfer from audio prompts and multi-level style control from natural language prompts, using blurred boundary encoders, contrastive prompt alignment, and a flow-based transformer.