REVIEW 3 major objections 5 minor 135 references
This paper claims that a single optimization recipe over a paired global-and-spatial latent representation turns any of several 3D generators into a semantic editor that follows copy, resize, delete, mix, drag, and region-drag operations wh
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 20:33 UTC pith:GUPA7HEM
load-bearing objection Useful category-agnostic 3D editing recipe with a clean operator-to-objective formulation, but the SOTA-across-six-operators claim rests on evidence for only two and an unmeasured locality premise. the 3 major comments →
CNS-Edit++: Category-Agnostic 3D Editing with Coupled Neural Shape Representation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that a coupled neural shape representation—a global latent code z plus a 3D neural feature volume F extracted from an intermediate generator layer—can be co-optimized to satisfy any editing operation. The objective L_op = |F_k[Γ] − sg(V)|_1 sets target feature values in a selected region; gradients flow back through the volume to update the global code, and the updated code is decoded to produce the edit. The paper argues that because the volume is taken from a layer followed only by local convolutions, local feature changes yield local shape changes, so the optimization steers the latent toward the user's instruction while keeping global semantics such as symmetry.
What carries the argument
The load-bearing mechanism is the pair (z, F): z encodes high-level shape meaning (from a shape encoder, refined text embedding, or image embedding), and F is a 3D feature volume sliced from a selected intermediate layer whose remaining processing has limited receptive field. Each operator defines coordinates Γ and target features V; the loss L_op plus preservation regularizer L_reg = ||(1−M)⊙(F_k−F_0)||_1 guides gradient steps on z, followed by re-extraction of F and re-evaluation, until the decoded shape satisfies the operation. The choice of layer depth is what ties locality to semantics.
Load-bearing premise
The whole scheme assumes feature values in the chosen intermediate layer are locally aligned with 3D space, so changing a feature at one coordinate changes the shape only near that location—and that gradients through the volume can reliably steer the global code to the requested edit.
What would settle it
Clamp a small feature patch to zero (or to copied features) while keeping the latent fixed, and measure where the geometry actually moves across layer depths and object categories. If changes spread across most of the shape, or if gradient-based latent updates fail to move the edit toward the target, the locality assumption is refuted.
If this is right
- Any pretrained 3D generator with an invertible latent path can be turned into an editor by adding the CNS volume extractor and the L_op objective; no retraining is needed.
- Operators are atomic and composable: copy plus delete yields a cut-paste operation, so users can construct new editing tasks from the six basic ones.
- Edits are semantic rather than purely geometric: resizing one part automatically updates symmetric counterparts.
- The method supports topology changes such as part removal or duplication, which pure deformation methods cannot.
- Unedited regions are preserved by combining a regularizer on the feature volume with cached key-value replacement during decoding.
Where Pith is reading between the lines
- The same recipe should transfer to video or scene generators that expose an intermediate feature volume, enabling drag-style editing in 4D—a direction the paper leaves to future work.
- Layer choice could be made principled by measuring coordinate alignment directly (perturb a feature patch and see where geometry moves) rather than relying on one backbone ablation.
- A stress test on categories far outside the generator's training distribution would reveal whether the locality premise holds in general; the reported benchmarks cover 20 Objaverse categories but no explicit out-of-distribution split.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CNS-Edit++, a latent-space 3D shape editing framework built on a Coupled Neural Shape (CNS) representation that couples a global latent code z with a 3D neural feature volume F. The representation is constructed on three backbones (a category-specific wavelet diffusion model [13], and the category-agnostic TRELLIS [14] and Direct3D-S2 [15] foundation models). Edits are posed as an objective L_op = |F_k[Γ]−sg(V)|_1 (Eq. 8) defined on coordinates Γ of the feature volume, together with a region preservation loss L_reg (Eq. 12), and the two components are co-optimized. Six operators are introduced: copy, resize, delete, mix, point-wise drag, and region-wise drag, alongside KV-cache replacement and latent feature regularization for preserving unedited regions. Quantitative comparisons are reported for point-wise drag (Table 2) and delete (Table 3) on small author-built benchmarks, with additional qualitative results for the remaining operators, and a category-specific comparison in Table 1.
Significance. The paper's central ambition — a single optimization recipe that turns different 3D generative backbones into spatially controllable editors — is valuable and timely. The formulation is clean and largely backbone-agnostic, the second-order inversion and KV-cache/regularization mechanisms are clearly motivated, and the ablations (Table 4, Figs. 22–23) confirm that each proposed component contributes. The qualitative results are broad and show that topology-changing edits are achievable. However, the central claim rests on a locality premise for the feature-volume-to-shape mapping that is asserted rather than directly measured, and the quantitative SOTA claim is supported for only two operators on 50-case benchmarks with no significance testing. If the mechanism holds, this would be an important step toward general-purpose 3D editing; the evidence as presented is not yet conclusive.
major comments (3)
- [Sec. 4.1, Eq. (8), Fig. 23] The entire editing mechanism rests on the claim that local modifications to the neural feature volume F correspond to localized shape changes ('local modifications in this volume typically correspond to targeted changes in the desired regions'). For the wavelet backbone this is supported by the local support property of wavelets, but for TRELLIS and Direct3D-S2, F is extracted from layer 21 of a rectified-flow transformer with no analogous argument. The only evidence is the layer-depth ablation in Fig. 23 (J=9 vs J=15 vs J=12), which shows layer choice matters but does not measure spatial locality. The paper's own word 'typically' flags the gap. I ask for a direct locality diagnostic: perturb a spatially small region of F, decode the shape, and quantify the change outside the target region; or measure the Jacobian of the decoded geometry with respect to F coordinates. This is load-bearin
- [Tables 2–3 and Sec. 5.2] Quantitative state-of-the-art comparisons are provided for only two of the six operators: point-wise drag and delete. Copy, resize, mix, and region-wise drag appear only in qualitative figures (Figs. 15–21). Given the abstract's claim of 'extensive quantitative and qualitative evaluations' and the conclusion's 'outperforming existing works,' the quantitative claim should be either extended to the other operators (e.g., with simple geometric metrics such as target-region overlap or feature matching) or explicitly scoped to the two evaluated operators. Without this, the 'six operators' contribution is supported only by qualitative evidence.
- [Sec. 5.3, Tables 1–3] The benchmarks use 50 editing cases per setting and a 10-participant user study, with no error bars, confidence intervals, or significance tests. FID/KID estimated on 50 samples are high-variance, especially when rendered from a small set of shapes; some reported gaps (e.g., Table 3: Ours Direct3D-S2 94.26 vs VoxHammer 100.37) may be within noise, while others are large. Please report variance across random seeds or bootstrap intervals, give per-case or per-category breakdowns, and state user-study inter-rater variability. This is necessary to support the 'consistent improvements' claim.
minor comments (5)
- [Sec. 5.1] The drag radii (r1, r2) = (1, 2) are given without units. Feature volumes from the three backbones have different resolutions; clarify whether these are voxel units and whether the same values are appropriate across backbones.
- [Table 4 caption] The caption does not state which operator and backbone the ablation is evaluated on. Please specify (it appears to be point-wise drag on TRELLIS-image, but this should be explicit).
- [Eqs. (3)–(7)] The second-order Taylor derivation is described as including a derivative term, but the final update is the standard midpoint rule x_t2 ≈ x_t1 + Δ·v_m. Clarify the relationship between the Taylor expansion and the midpoint update to avoid confusion.
- [Eq. (9)] Point tracking uses L1 feature distance on raw features. A brief comment on why this is preferred over cosine similarity or other metrics, and on failure modes when tracking drifts, would strengthen the presentation.
- [General] There is no direct comparison with the authors' previous CNS-Edit in the category-agnostic setting. Since the '++' contribution is the generalization, a comparison against a CNS-Edit variant adapted to foundation models (or an explicit explanation of why such comparison is omitted) would clarify the incremental benefit.
Circularity Check
No significant circularity: the editing results are obtained by optimization with externally evaluated outcomes; operator objectives are not self-fulfilling predictions.
full rationale
The paper's claimed derivation is an optimization recipe (Sec. 4.3), not a fitted prediction. The operator objectives (Eq. 8) specify desired feature-volume values, but the edited shape is obtained by decoding the co-optimized latent z_N (Sec. 4.3), and success is measured externally: user-study QS/MS scores and FID/KID against DeepMetaHandles, SLIDE, DualSDF, SPAGHETTI, APAP, MVDrag3D, DragFlow, Tailor3D, Instant3dit, and VoxHammer (Tables 1-3). No parameter is fit to the target edit and then reported as the output of a prediction. The locality premise that neural feature volumes are spatially aligned with shape regions (Sec. 4.1) is an empirical assumption, supported for the wavelet backbone by the local-support argument and for all backbones by a layer-depth ablation (Fig. 23, Table 4); it is not made true by definition or by a self-citation. The authors cite their own previous wavelet-diffusion backbone [13], [74] and CLIPXplore [12], but the category-agnostic claims are instantiated on external TRELLIS [14] and Direct3D-S2 [15], so these self-citations are not load-bearing. The unverified nature of the feature-volume locality premise is a correctness risk, not a circularity: it is a stated assumption about the learned generator's intermediate features, not an identity built into the objective.
Axiom & Free-Parameter Ledger
free parameters (6)
- Feature-layer index for volume F (J=12 for [13]'s 16-layer U-Net; J=21 for TRELLIS and Direct3D-S2) =
12 / 21 / 21
- Feature-extraction timestep t =
200 (diffusion inversion backbone); 0.7 (TRELLIS, Direct3D-S2)
- Region-preservation weight λ_reg (Eq. 13) =
not stated in main text
- CNS co-optimization hyperparameters =
Adam lr 3e-2 / 2e-3 / 2e-3 per backbone; N max iterations and K inversion steps unspecified
- Drag geometry parameters =
r1=1, r2=2 (voxel units)
- Termination thresholds =
loss < 1/3 of initial; drag stops when source reaches target
axioms (6)
- domain assumption Intermediate features of the pre-trained backbones (U-Net layer 12; TRELLIS/Direct3D-S2 layer 21) are locally aligned with 3D space, so that local feature edits map to local shape edits.
- domain assumption The co-optimization loop (gradient on z through the differentiable volume extraction) converges to a code z_N that decodes to a plausible shape satisfying the edit.
- ad hoc to paper For text-conditioned backbones, optimizing the CLIP text embedding with the rectified-flow objective (Eq. 1) yields an instance-specific code inside the generator's support.
- domain assumption The second-order Taylor update with midpoint velocity (Eqs. 3–7) approximates the rectified-flow inversion trajectory well enough for faithful feature-volume construction.
- domain assumption A user-provided binary mask M is available and can be mapped into neural-volume/token space for L_reg and KV-cache replacement.
- domain assumption FID/KID on rendered views and 10-participant QS/MS ratings are valid proxies for editing fidelity and user-operation fulfillment.
invented entities (1)
-
Coupled Neural Shape (CNS) representation — global latent code z + 3D neural feature volume F
no independent evidence
read the original abstract
This paper presents a latent-space 3D shape editing framework built upon a coupled neural shape (CNS) representation and a neural feature volume optimization. This work extends CNS-Edit, built on Coupled Neural Shape optimization, to CNS-Edit++, by generalizing the category-specific coupled representation to category-agnostic 3D shape editing with foundation models. The Coupled Neural Shape (CNS) representation couples a global latent code that captures high-level shape semantics with a 3D neural feature volume that provides spatial context for local shape manipulation. Then we formulate a coupled neural shape optimization procedure that co-optimizes these two components subject to a given editing operation. Our framework can be instantiated on both the category-specific 3D inversion model and category-agnostic 3D foundation models. We provide various shape editing operators, including copy, resize, delete, mix, point-wise drag, and region-wise drag, each of which is formulated as an objective to guide the CNS optimization. To preserve regions outside the editing area, we further introduce two complementary region-wise control mechanisms, i.e., KV-cache replacement and latent feature regularization. Extensive quantitative and qualitative evaluations across different 3D generative models demonstrate the strong capabilities of our approach over state-of-the-art solutions.
Figures
Reference graph
Works this paper leans on
-
[1]
A revisit of shape editing techniques: From the geometric to the neural viewpoint,
Y.-J. Yuan, Y.-K. Lai, T. Wu, L. Gao, and L. Liu, “A revisit of shape editing techniques: From the geometric to the neural viewpoint,” Journal of Computer Science and Technology, vol. 36, no. 3, pp. 520– 554, 2021
2021
-
[2]
Generative adver- sarial nets,
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde- Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adver- sarial nets,” inConference on Neural Information Processing Systems (NeurIPS), 2014, pp. 2672–2680
2014
-
[3]
Denoising diffusion probabilis- tic models,
J. Ho, A. Jain, and P . Abbeel, “Denoising diffusion probabilis- tic models,”Conference on Neural Information Processing Systems (NeurIPS), pp. 6840–6851, 2020
2020
-
[4]
Drag your GAN: Interactive point-based manip- ulation on the generative image manifold,
X. Pan, A. Tewari, T. Leimk ¨uhler, L. Liu, A. Meka, and C. Theobalt, “Drag your GAN: Interactive point-based manip- ulation on the generative image manifold,” inProceedings of SIGGRAPH, 2023, pp. 1–11
2023
-
[5]
Dragdif- fusion: Harnessing diffusion models for interactive point-based image editing,
Y. Shi, C. Xue, J. Pan, W. Zhang, V . Y. Tan, and S. Bai, “Dragdif- fusion: Harnessing diffusion models for interactive point-based image editing,”arXiv preprint arXiv:2306.14435, 2023
Pith/arXiv arXiv 2023
-
[6]
Dragondif- fusion: Enabling drag-style manipulation on diffusion models,
C. Mou, X. Wang, J. Song, Y. Shan, and J. Zhang, “Dragondif- fusion: Enabling drag-style manipulation on diffusion models,” arXiv preprint arXiv:2307.02421, 2023
Pith/arXiv arXiv 2023
-
[7]
Dragvideo: Interactive drag-style video editing,
Y. Deng, R. Wang, Y. Zhang, Y.-W. Tai, and C.-K. Tang, “Dragvideo: Interactive drag-style video editing,”arXiv preprint arXiv:2312.02216, 2023
Pith/arXiv arXiv 2023
-
[8]
Pointgmm: A neural gmm network for point clouds,
A. Hertz, R. Hanocka, R. Giryes, and D. Cohen-Or, “Pointgmm: A neural gmm network for point clouds,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 12 054– 12 063
2020
-
[9]
DualSDF: Semantic shape manipulation using a two-level representation,
Z. Hao, H. Averbuch-Elor, N. Snavely, and S. Belongie, “DualSDF: Semantic shape manipulation using a two-level representation,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 7631–7641
2020
-
[10]
Neural template: Topology-aware reconstruction and disentangled gen- eration of 3d meshes,
K.-H. Hui*, R. Li*, J. Hu, and C.-W. F. . joint first authors), “Neural template: Topology-aware reconstruction and disentangled gen- eration of 3d meshes,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 18 572–18 582
2022
-
[11]
Salad: Part-level latent diffusion for 3d shape generation and manipulation,
J. Koo, S. Yoo, M. H. Nguyen, and M. Sung, “Salad: Part-level latent diffusion for 3d shape generation and manipulation,” in IEEE International Conference on Computer Vision (ICCV), 2023, pp. 14 441–14 451
2023
-
[12]
Clipxplore: Coupled clip and shape spaces for 3d shape exploration,
J. Hu*, K.-H. Hui*, Z. Liu, H. Zhang, and C.-W. Fu, “Clipxplore: Coupled clip and shape spaces for 3d shape exploration,” in Proceedings of SIGGRAPH Asia, 2023, pp. 1–12
2023
-
[13]
Neural wavelet- domain diffusion for 3d shape generation, inversion, and manip- ulation,
J. Hu*, K.-H. Hui*, Z. Liu, R. Li, and C.-W. Fu, “Neural wavelet- domain diffusion for 3d shape generation, inversion, and manip- ulation,”ACM Transactions on Graphics (TOG), 2023
2023
-
[14]
Structured 3d latents for scalable and versatile 3d generation,
J. Xiang, Z. Lv, S. Xu, Y. Deng, R. Wang, B. Zhang, D. Chen, X. Tong, and J. Yang, “Structured 3d latents for scalable and versatile 3d generation,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 1–10
2025
-
[15]
Direct3d-s2: Gigascale 3d generation made easy with spatial sparse attention,
S. Wu, Y. Lin, F. Zhang, Y. Zeng, Y. Yang, Y. Bao, J. Qian, S. Zhu, X. Cao, P . Torret al., “Direct3d-s2: Gigascale 3d generation made easy with spatial sparse attention,” inConference on Neural Information Processing Systems (NeurIPS), 2025
2025
-
[16]
As-rigid-as-possible surface model- ing,
O. Sorkine and M. Alexa, “As-rigid-as-possible surface model- ing,” inSymposium on Geometry processing, vol. 4, 2007, pp. 109– 116
2007
-
[17]
Har- monic coordinates for character articulation,
P . Joshi, M. Meyer, T. DeRose, B. Green, and T. Sanocki, “Har- monic coordinates for character articulation,”ACM Transactions on Graphics, vol. 26, no. 3, pp. 71–es, 2007
2007
-
[18]
Mean value coordinates for closed triangular meshes,
T. Ju, S. Schaefer, and J. Warren, “Mean value coordinates for closed triangular meshes,”ACM Transactions on Graphics, vol. 24, no. 3, p. 561–566, 2005
2005
-
[19]
Green coordinates,
Y. Lipman, D. Levin, and D. Cohen-Or, “Green coordinates,” ACM Transactions on Graphics, vol. 27, no. 3, pp. 1–10, 2008
2008
-
[20]
Green coordinates for triquad cages in 3d,
J.-M. Thiery and T. Boubekeur, “Green coordinates for triquad cages in 3d,” inProceedings of SIGGRAPH Asia, 2022, pp. 1–8
2022
-
[21]
Differential coordinates for interactive mesh editing,
Y. Lipman, O. Sorkine, D. Cohen-Or, D. Levin, C. Rossi, and H.-P . Seidel, “Differential coordinates for interactive mesh editing,” in Proceedings of IEEE International Conference on Shape Modeling and Applications, 2004, pp. 181–190
2004
-
[22]
Slippage-preserving reshaping of human-made 3d content,
C. Ara ´ujo, N. Vining, S. Burla, M. Ruivo de Oliveira, E. Rosales, and A. Sheffer, “Slippage-preserving reshaping of human-made 3d content,”ACM Transactions on Graphics (SIGGRAPH Asia), vol. 42, no. 6, 2023. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 16
2023
-
[23]
Laplacian surface editing,
O. Sorkine, D. Cohen-Or, Y. Lipman, M. Alexa, C. R ¨ossl, and H.-P . Seidel, “Laplacian surface editing,” inEurographics Symposium on Geometry Processing (SGP), 2004, pp. 175–184
2004
-
[24]
Structure-aware shape processing,
N. Mitra, M. Wand, H. Zhang, D. Cohen-Or, and M. Bokeloh, “Structure-aware shape processing,” inEurographics State-of-the- art Report (STAR), 2013
2013
-
[25]
iWIRES: an analyze-and-edit approach to shape manipulation,
R. Gal, O. Sorkine, N. J. Mitra, and D. Cohen-Or, “iWIRES: an analyze-and-edit approach to shape manipulation,”ACM Trans- actions on Graphics (SIGGRAPH), 2009
2009
-
[26]
Component-wise controllers for structure-preserving shape ma- nipulation,
Y. Zheng, H. Fu, D. Cohen-Or, O. K.-C. Au, and C.-L. Tai, “Component-wise controllers for structure-preserving shape ma- nipulation,”Computer Graphics Forum (Eurographics), 2011
2011
-
[27]
Symmetry hierarchy of man-made objects,
Y. Wang, K. Xu, J. Li, H. Zhang, A. Shamir, L. Liu, Z. Cheng, and Y. Xiong, “Symmetry hierarchy of man-made objects,”Computer Graphics Forum (Eurographics), vol. 30, no. 2, pp. 287–296, 2011
2011
-
[28]
COALESCE: Component assembly by learning to synthesize connections,
K. Yin, Z. Chen, S. Chaudhuri, M. Fisher, V . Kim, and H. Zhang, “COALESCE: Component assembly by learning to synthesize connections,” inProc. of 3DV, 2020
2020
-
[29]
Dag amendment for inverse control of parametric shapes,
E. Michel and T. Boubekeur, “Dag amendment for inverse control of parametric shapes,”ACM Transactions on Graphics (TOG), vol. 40, no. 4, pp. 1–14, 2021
2021
-
[30]
Differentiable 3d cad programs for bidirectional editing,
D. Cascaval, M. Shalah, P . Quinn, R. Bodik, M. Agrawala, and A. Schulz, “Differentiable 3d cad programs for bidirectional editing,” inComputer Graphics Forum, vol. 41, no. 2, 2022, pp. 309–323
2022
-
[31]
Neuform: Adaptive overfitting for neural shape editing,
C. Z. Lin, N. J. Mitra, G. Wetzstein, L. Guibas, and P . Guerrero, “Neuform: Adaptive overfitting for neural shape editing,” in Conference on Neural Information Processing Systems (NeurIPS), 2022
2022
-
[32]
Generating part-aware editable 3d shapes without 3d supervision,
K. Tertikas, D. Paschalidou, B. Pan, J. J. Park, M. A. Uy, I. Emiris, Y. Avrithis, and L. Guibas, “Generating part-aware editable 3d shapes without 3d supervision,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 4466–4478
2023
-
[33]
3dn: 3d deformation network,
W. Wang, D. Ceylan, R. Mech, and U. Neumann, “3dn: 3d deformation network,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 1038–1046
2019
-
[34]
Neural cages for detail-preserving 3d deformations,
W. Yifan, N. Aigerman, V . G. Kim, S. Chaudhuri, and O. Sorkine- Hornung, “Neural cages for detail-preserving 3d deformations,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 75–83
2020
-
[35]
Deepmetahandles: Learn- ing deformation meta-handles of 3d meshes with biharmonic coordinates,
M. Liu, M. Sung, R. Mech, and H. Su, “Deepmetahandles: Learn- ing deformation meta-handles of 3d meshes with biharmonic coordinates,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 12–21
2021
-
[36]
Shapeflow: Learnable deformation flows among 3d shapes,
C. Jiang, J. Huang, A. Tagliasacchi, and L. J. Guibas, “Shapeflow: Learnable deformation flows among 3d shapes,” inConference on Neural Information Processing Systems (NeurIPS), 2020, pp. 9745– 9757
2020
-
[37]
Mesh draping: Parametrization-free neural mesh transfer,
A. Hertz, O. Perel, R. Giryes, O. Sorkine-Hornung, and D. Cohen- Or, “Mesh draping: Parametrization-free neural mesh transfer,” inComputer Graphics Forum, vol. 42, no. 1, 2023, pp. 72–85
2023
-
[38]
Neural shape deformation priors,
J. Tang, L. Markhasin, B. Wang, J. Thies, and M. Nießner, “Neural shape deformation priors,” inConference on Neural Information Processing Systems (NeurIPS), 2022, pp. 17 117–17 132
2022
-
[39]
Exim: A hybrid explicit-implicit representation for text-guided 3d shape generation,
Z. Liu, J. Hu, K.-H. Hui, X. Qi, D. Cohen-Or, and C.-W. Fu, “Exim: A hybrid explicit-implicit representation for text-guided 3d shape generation,”ACM Transactions on Graphics (SIGGRAPH Asia), vol. 42, no. 6, pp. 1–12, 2023
2023
-
[40]
Shapecrafter: A recursive text-conditioned 3d shape generation model,
R. Fu, X. Zhan, Y. Chen, D. Ritchie, and S. Sridhar, “Shapecrafter: A recursive text-conditioned 3d shape generation model,” in Conference on Neural Information Processing Systems (NeurIPS), 2022
2022
-
[41]
Towards implicit text- guided 3d shape generation,
Z. Liu, Y. Wang, X. Qi, and C.-W. Fu, “Towards implicit text- guided 3d shape generation,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 17 896–17 906
2022
-
[42]
Changeit3D: Language-Assisted 3D Shape Edits and Deforma- tions,
P . Achlioptas, I. Huang, M. Sung, S. Tulyakov, and L. Guibas, “Changeit3D: Language-Assisted 3D Shape Edits and Deforma- tions,” inIEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), 2023
2023
-
[43]
Ladis: Language disentanglement for 3d shape edit- ing,
I. Huang, P . Achlioptas, T. Zhang, S. Tulyakov, M. Sung, and L. Guibas, “Ladis: Language disentanglement for 3d shape edit- ing,” inIn Findings of Empirical Methods in Natural Language Processing (EMNLP), 2022
2022
-
[44]
Instruct-nerf2nerf: Editing 3d scenes with instructions,
A. Haque, M. Tancik, A. A. Efros, A. Holynski, and A. Kanazawa, “Instruct-nerf2nerf: Editing 3d scenes with instructions,” inIEEE International Conference on Computer Vision (ICCV), 2023, pp. 19 740–19 750
2023
-
[45]
Voxhammer: Training-free precise and coherent 3d editing in native 3d space,
L. Li, Z. Huang, H. Feng, G. Zhuang, R. Chen, C. Guo, and L. Sheng, “Voxhammer: Training-free precise and coherent 3d editing in native 3d space,” inInternational Conference on 3D Vision (3DV), 2026, pp. 1281–1292
2026
-
[46]
Tailor3d: Customized 3d assets editing and generation with dual-side images,
Z. Qi, Y. Yang, M. Zhang, L. Xing, X. Wu, T. Wu, D. Lin, X. Liu, J. Wang, and H. Zhao, “Tailor3d: Customized 3d assets editing and generation with dual-side images,”arXiv preprint arXiv:2407.06191, 2024
Pith/arXiv arXiv 2024
-
[47]
Instant3dit: Multiview inpainting for fast edit- ing of 3d objects,
A. Barda, M. Gadelha, V . G. Kim, N. Aigerman, A. H. Bermano, and T. Groueix, “Instant3dit: Multiview inpainting for fast edit- ing of 3d objects,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 16 273–16 282
2025
-
[48]
Mamba as a bridge: Where vi- sion foundation models meet vision language models for domain-generalized semantic segmentation,
X. Zhang and R. T. Tan, “Mamba as a bridge: Where vi- sion foundation models meet vision language models for domain-generalized semantic segmentation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 14 527–14 537
2025
-
[49]
Pegasus: 3d personalization of geometry and appearance,
J. Hu, B. Hu, K.-H. Hui, H. Li, Z. Liu, D. Cohen-Or, and C.-W. Fu, “Pegasus: 3d personalization of geometry and appearance,” inProceedings of SIGGRAPH, 2026
2026
-
[50]
Hierarchical neural semantic representation for 3d semantic correspondence,
K. Du, J. Hu, H. Li, H. Xu, H. Huang, C.-W. Fu, and S. Liu, “Hierarchical neural semantic representation for 3d semantic correspondence,” inProceedings of SIGGRAPH Asia, 2025, pp. 1– 11
2025
-
[51]
Sketchsampler: Sketch-based 3d reconstruction via view-dependent depth sam- pling,
C. Gao, Q. Yu, L. Sheng, Y.-Z. Song, and D. Xu, “Sketchsampler: Sketch-based 3d reconstruction via view-dependent depth sam- pling,” inEuropean Conference on Computer Vision (ECCV), 2022, pp. 464–479
2022
-
[52]
Sketch2model: View- aware 3d modeling from single free-hand sketches,
S.-H. Zhang, Y.-C. Guo, and Q.-W. Gu, “Sketch2model: View- aware 3d modeling from single free-hand sketches,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 6012–6021
2021
-
[53]
Sketch2mesh: Reconstructing and editing 3d shapes from sketches,
B. Guillard, E. Remelli, P . Yvernay, and P . Fua, “Sketch2mesh: Reconstructing and editing 3d shapes from sketches,” inIEEE In- ternational Conference on Computer Vision (ICCV), 2021, pp. 13 023– 13 032
2021
-
[54]
Locally attentional sdf diffusion for controllable 3d shape generation,
X.-Y. Zheng, H. Pan, P .-S. Wang, X. Tong, Y. Liu, and H.-Y. Shum, “Locally attentional sdf diffusion for controllable 3d shape generation,”ACM Transactions on Graphics (SIGGRAPH), vol. 42, no. 4, 2023
2023
-
[55]
Sked: Sketch-guided text-based 3d editing,
A. Mikaeili, O. Perel, M. Safaee, D. Cohen-Or, and A. Mahdavi- Amiri, “Sked: Sketch-guided text-based 3d editing,” inIEEE In- ternational Conference on Computer Vision (ICCV), 2023, pp. 14 607– 14 619
2023
-
[56]
Learn- ing a probabilistic latent space of object shapes via 3D generative- adversarial modeling,
J. Wu, C. Zhang, T. Xue, B. Freeman, and J. Tenenbaum, “Learn- ing a probabilistic latent space of object shapes via 3D generative- adversarial modeling,” inConference on Neural Information Process- ing Systems (NeurIPS), 2016, pp. 82–90
2016
-
[57]
Improved adversarial systems for 3D object generation and reconstruction,
E. J. Smith and D. Meger, “Improved adversarial systems for 3D object generation and reconstruction,” inConference on Robot Learning, 2017, pp. 87–96
2017
-
[58]
3D shape generation and completion through point-voxel diffusion,
L. Zhou, Y. Du, and J. Wu, “3D shape generation and completion through point-voxel diffusion,” inIEEE International Conference on Computer Vision (ICCV), 2021, pp. 5826–5835
2021
-
[59]
Diffusion probabilistic models for 3D point cloud generation,
S. Luo and W. Hu, “Diffusion probabilistic models for 3D point cloud generation,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 2837–2845
2021
-
[60]
Lion: Latent point diffusion models for 3d shape generation,
X. Zeng, A. Vahdat, F. Williams, Z. Gojcic, O. Litany, S. Fidler, and K. Kreis, “Lion: Latent point diffusion models for 3d shape generation,” inConference on Neural Information Processing Systems (NeurIPS), 2022
2022
-
[61]
Point-e: A system for generating 3d point clouds from complex prompts,
A. Nichol, H. Jun, P . Dhariwal, P . Mishkin, and M. Chen, “Point-e: A system for generating 3d point clouds from complex prompts,” arXiv preprint arXiv:2212.08751, 2022
Pith/arXiv arXiv 2022
-
[62]
SP-GAN: Sphere-guided 3D shape generation and manipulation,
R. Li, X. Li, K.-H. Hui, and C.-W. Fu, “SP-GAN: Sphere-guided 3D shape generation and manipulation,”ACM Transactions on Graphics (SIGGRAPH), vol. 40, no. 4, 2021
2021
-
[63]
MRGAN: Multi-rooted 3D shape generation with unsupervised part disen- tanglement,
R. Gal, A. Bermano, H. Zhang, and D. Cohen-Or, “MRGAN: Multi-rooted 3D shape generation with unsupervised part disen- tanglement,” inIn ICCV Workshop on Structural and Compositional Learning on 3D Data (StruCo3D)., 2020, pp. 2039–2048
2020
-
[64]
Progressive point cloud deconvolution generation network,
L. Hui, R. Xu, J. Xie, J. Qian, and J. Yang, “Progressive point cloud deconvolution generation network,” inEuropean Conference on Computer Vision (ECCV), 2020, pp. 397–413
2020
-
[65]
Geocomplete: Geometry-aware diffusion for reference-driven image completion,
B. Lin, T. Chen, and R. Tan, “Geocomplete: Geometry-aware diffusion for reference-driven image completion,”Advances in JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 17 Neural Information Processing Systems, vol. 38, pp. 35 238–35 259, 2026
2021
-
[66]
Glowgs: Generative seman- tic feature learning for 3d gaussian splatting in nighttime glow scenes,
B. Lin, X. Cao, J. Guo, and R. T. Tan, “Glowgs: Generative seman- tic feature learning for 3d gaussian splatting in nighttime glow scenes,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026, pp. 275–284
2026
-
[67]
Raindropgs: A benchmark for 3d gaussian splatting under rain- drop conditions,
Z. Teng, T. Chen, B. Lin, Z. Yuan, X. Li, X. Zhang, and S. Zhang, “Raindropgs: A benchmark for 3d gaussian splatting under rain- drop conditions,”arXiv preprint arXiv:2510.17719, 2025
arXiv 2025
-
[68]
Adafit: Rethinking learning-based normal estimation on point clouds,
R. Zhu, Y. Liu, Z. Dong, Y. Wang, T. Jiang, W. Wang, and B. Yang, “Adafit: Rethinking learning-based normal estimation on point clouds,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 6118–6127
2021
-
[69]
MeshDiffusion: Score-based generative 3D mesh model- ing,
Z. Liu, Y. Feng, M. J. Black, D. Nowrouzezahrai, L. Paull, and W. Liu, “MeshDiffusion: Score-based generative 3D mesh model- ing,” inInternational Conference on Learning Representations, 2023
2023
-
[70]
Meshgpt: Generating tri- angle meshes with decoder-only transformers,
Y. Siddiqui, A. Alliegro, A. Artemov, T. Tommasi, D. Sirigatti, V . Rosov, A. Dai, and M. Nießner, “Meshgpt: Generating tri- angle meshes with decoder-only transformers,”arXiv preprint arXiv:2311.15475, 2023
Pith/arXiv arXiv 2023
-
[71]
Bsp-net: Generating compact meshes via binary space partitioning,
Z. Chen, A. Tagliasacchi, and H. Zhang, “Bsp-net: Generating compact meshes via binary space partitioning,” inIEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 45–54
2020
-
[72]
Las- comp: Zero-shot 3d completion with latent-spatial consistency,
W. Yan, H. Li, H. Xu, N. Ye, Y. Ai, S. Liu, and J. Hu, “Las- comp: Zero-shot 3d completion with latent-spatial consistency,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2026, pp. 7588–7599
2026
-
[73]
4dpc2hat: Towards dynamic point cloud understanding with failure-aware bootstrapping,
X. Zhang, W. Yan, Y. Shi, X. Qiu, T. He, Y. Li, M. Li, and H. Fan, “4dpc2hat: Towards dynamic point cloud understanding with failure-aware bootstrapping,”arXiv preprint arXiv:2602.03890, 2026
Pith/arXiv arXiv 2026
-
[74]
Neural wavelet-domain diffusion for 3d shape generation,
K.-H. Hui, R. Li, J. Hu, and C.-W. Fu, “Neural wavelet-domain diffusion for 3d shape generation,” inProceedings of SIGGRAPH Asia, 2022, pp. 1–9
2022
-
[75]
Get3d: A generative model of high quality 3d textured shapes learned from images,
J. Gao, T. Shen, Z. Wang, W. Chen, K. Yin, D. Li, O. Litany, Z. Gojcic, and S. Fidler, “Get3d: A generative model of high quality 3d textured shapes learned from images,” inConference on Neural Information Processing Systems (NeurIPS), 2022
2022
-
[76]
3dshape2vecset: A 3d shape representation for neural fields and generative diffusion models,
B. Zhang, J. Tang, M. Nießner, and P . Wonka, “3dshape2vecset: A 3d shape representation for neural fields and generative diffusion models,”ACM Transactions on Graphics (SIGGRAPH), vol. 42, no. 4, 2023
2023
-
[77]
Occupancy networks: Learning 3D reconstruction in function space,
L. Mescheder, M. Oechsle, M. Niemeyer, S. Nowozin, and A. Geiger, “Occupancy networks: Learning 3D reconstruction in function space,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 4460–4470
2019
-
[78]
Learning implicit fields for generative shape modeling,
Z. Chen and H. Zhang, “Learning implicit fields for generative shape modeling,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 5939–5948
2019
-
[79]
3D shape generation with grid- based implicit functions,
M. Ibing, I. Lim, and L. Kobbelt, “3D shape generation with grid- based implicit functions,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 13 559–13 568
2021
-
[80]
Ssp: Semi-signed prioritized neural fitting for surface reconstruction from unoriented point clouds,
R. Zhu, D. Kang, K.-H. Hui, Y. Qian, S. Qiu, Z. Dong, L. Bao, P .-A. Heng, and C.-W. Fu, “Ssp: Semi-signed prioritized neural fitting for surface reconstruction from unoriented point clouds,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 3769–3778
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.