REVIEW 4 major objections 6 minor 52 references
Position-aware Guided Point Cloud Completion with CLIP Model
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper claims that attaching a frozen CLIP branch and a 24-parameter position-aware module to a point cloud completion network improves completion accuracy across multiple baselines.
desk verdict Incremental CLIP-based point cloud completion with tiny gains and a position-aware module whose mechanism is asserted but unsupported; worth a referee but needs significant revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two pieces carry the argument. The CLIP-enhanced module turns each partial point cloud into a Point-Text-Image triplet, using a category template sentence and six orthogonal depth projections processed by six shared frozen CLIP encoders, so the network inherits pretrained semantic alignment at no training cost. The position-aware module is the spatial mechanism: it splits each projection into 2x2 non-overlapping blocks, learns one weight per block per projection (24 scalars total) with unselected blocks held at 1 during training, and cross-attends these local features with inpainted global projection features to produce a fusion feature that guides the decoder. The paper treats these weights as marking which block contains the projection of the missing part.
What would settle it
Compare the full model against the same model with the 24 block weights replaced by fixed constants, or randomly permuted per input; if completion quality does not drop materially, the weights are not carrying location information. A direct check is to visualize the learned weights for shapes with known missing sides and see whether the high-weight block consistently matches the missing side.
Extended reading notes
Core claim
The paper's central claim is that point cloud completion improves when the decoder receives both CLIP-extracted semantic features and learned indicators of where the missing geometry lies. The method projects the partial point cloud onto six coordinate planes, pairs those projections with a simple category sentence, and runs them through shared frozen CLIP encoders; an EdgeConnect-inpainted version of each projection supplies global structure, while a position-aware module learns 24 block weights across the six 2x2 grids. Those local and global image features are fused and concatenated with point-cloud features before the coarse-to-fine decoder produces the complete point cloud. The authors state that this outperforms state-of-the-art completion methods on PCN and MVP and transfers to LiDAR scans on KITTI, with the main gains credited to the CLIP-enhanced module and the position-aware module.
Load-bearing premise
The load-bearing premise is that the 24 learned block weights encode where the missing part is located, but the paper does not show this; if they only act as generic scale factors, the position-aware mechanism collapses and the small gains could come from the CLIP features or extra capacity alone.
Editorial extensions
If this is right
- The same extension can be applied to any unimodal point cloud completion baseline; the paper reports gains on SnowflakeNet, PoinTr, AdaPoinTr, and CRA-PCN.
- A multimodal Point-Text-Image corpus can be built automatically from a unimodal dataset, since the text is a category template and the images are computed projections.
- Removing the position-aware module lowers performance, which the paper attributes to a loss of missing-location information.
- Because CLIP is frozen, the added computation is mostly in fusion layers and six forward passes through a shared image encoder, matching the paper's claim of a rapid and efficient upgrade.
Reading between the lines
- The paper never visualizes or ablates the 24 learned weights, so the most direct test of its spatial story is to inspect whether high-weight blocks track the side of the shape that is actually missing.
- If those weights behave as generic per-block gain factors rather than location indicators, the reported gains could still come from CLIP semantic features or added capacity, and a control experiment with randomized or fixed weights would separate those explanations.
- The same projection-plus-template-text setup could be reused for other point cloud tasks, such as part segmentation or missing-region detection, where CLIP's category-level alignment may substitute for part-level labels.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a multimodal extension for point cloud completion. Starting from a unimodal baseline (CRA-PCN by default), the method adds a frozen CLIP branch that consumes six orthographic projection images of the partial point cloud plus a category-level text template, and a 'position-aware module' that learns 24 scalar weights for 2x2 blocks across the six projections. The authors also construct PCN-TI and MVP-TI multimodal corpora and report experiments on PCN, MVP, and KITTI, with average L1 CD improving from 6.39 to 6.34 on PCN and L2 CD from 5.33 to 5.32 on MVP. The central claims are that CLIP text/image features improve completion and that the learned block weights localize the missing parts.
Significance. If the mechanism were established, the paper would offer a lightweight recipe for upgrading unimodal point cloud completion datasets with CLIP-based multimodal supervision, and the PCN-TI/MVP-TI corpora could be a useful community resource. The use of a frozen CLIP encoder and the simple template-based text generation are practical strengths. However, the reported gains are very small, no statistical support is provided, and the paper does not demonstrate that the position-aware module actually encodes per-instance missing-region information. As it stands, the evidence supports an incremental engineering contribution rather than the strong state-of-the-art and localization claims made in the abstract and methodology.
major comments (4)
- [Methodology, Position-aware Module] The claimed mechanism is not supported by the implementation as described. The 24 block weights are global learned scalars shared across all training and test instances; because the missing region changes from instance to instance, fixed per-block constants cannot represent the instance-specific 'projection location of the missing part.' No visualization of the weights, no per-block accuracy analysis, and no permutation test is provided. The only quantitative support is Table 4, where adding this module changes CRA-PCN from 6.37 to 6.34 (a difference of 0.03 x 10^-3 in L1 CD), which could equally be attributed to extra capacity or run-to-run noise. The central claim that the position-aware module localizes missing parts is therefore not established.
- [Tables 1, 3, and 4] The reported improvements are very small and no statistical evidence is given. The PCN average changes from 6.39 to 6.34, the MVP average from 5.33 to 5.32 with F1-Score unchanged at 0.529, and the position-aware ablation changes the average by 0.03 x 10^-3. Without error bars, multiple seeds, or significance tests, these differences are within plausible run-to-run variance. In addition, Table 1 shows that Ours is worse than CRA-PCN on Airplane (3.64 vs 3.59) and tied on Table, so the statement that the method 'achieves the best performance on 7/8 categories' is factually incorrect, and the abstract's claim that the method 'outperforms state-of-the-art point cloud completion methods' is overstated relative to the data.
- [Methodology, CLIP-enhanced Module; Experiments, Ablation Study] The contribution of the text modality is never isolated. The template 'There is {category} point cloud projection map' encodes only the category label, which is already available from the point cloud itself, and the ablations in Table 4 compare only with-CE versus with-CE+PA. There is no condition without text, without images, or with a category-only text condition. Consequently, the paper does not support the claim that CLIP text features provide 'richer detail information' beyond the class prior; the observed gains could come entirely from the projection-image branch.
- [Experiments, Ablation Study (Effectiveness of inpainting global projection map)] The CRA-PCN-GT row in Table 4 replaces the inpainted projection map with the projection map of the complete point cloud, which is an oracle condition that uses ground-truth information. The resulting improvement to 6.11 shows an upper bound, not the effectiveness of the proposed inpainting step. The paper never ablates the actual EdgeConnect-based inpainting against a simpler alternative, so the statement that 'the complete projection map provides more accurate positional information' is not supported as a property of the proposed pipeline.
minor comments (6)
- [Experiments, Results on KITTI; Table 2] The KITTI comparison protocol is unclear: the table caption says the baseline is AdaPoinTr, but AdaPoinTr is not listed, and there is no description of how MMD is computed on sparse LiDAR scans or how the comparison with SVDFormer is made fair.
- [Figure 3 and Methodology, Position-aware Module] The text examples in Figure 3 ('This is a map of car', 'There is a map of car missing a piece from the left side') are not the same templates as the ones described in the methodology; the inconsistency should be resolved.
- [Experiments, Ablation Study] The text refers to 'PointTr' instead of 'PoinTr' in the sentence 'PointTr and AdaPoinTr redefine the point cloud completion task'; please correct the typo.
- [Conclusion] The phrase 'loca-global weighted map' appears to be a typo for 'local-global weighted map'.
- [Datasets and Evaluation Metrics] The list of MVP baselines in Section Experiments includes 'CRN' twice; one of the entries should be removed or renamed.
- [General] No code, trained models, or dataset release links are provided; releasing the PCN-TI and MVP-TI corpora would make the contribution reproducible and easier to evaluate.
Circularity Check
No circularity: the CLIP-enhanced and position-aware modules are trained end-to-end and evaluated on external benchmarks, with no prediction reducing to a fitted constant or to a self-citation chain.
full rationale
The paper's derivation chain is self-contained with respect to circularity. The 24 block-weight parameters in the Position-aware Module are ordinary network parameters optimized by the point cloud completion loss; the paper does not fit them to the evaluation metric and then report a closely related quantity as a prediction. No equation in the paper defines the completed output in terms of a fitted constant in a way that would make the benchmark result an identity. The text descriptions are generated from the input point cloud's category label, and the six projection images are derived from the input point cloud itself, so the CLIP features are input-conditioned auxiliary signals rather than functions of the ground-truth completion target. The ablation labeled CRA-PCN-GT, which replaces the projection map with the complete point cloud's projection map, is explicitly an oracle comparison and not the proposed test-time pipeline, so it does not make the reported results circular. The only likely self-citation, Li (2023), appears in a general related-work list and is not load-bearing. The conclusion's stated limitation that position information and text description are not jointly modeled is an acknowledged weakness, not a circular step. The paper's interpretation that the learned 2x2 block weights 'ascertain if the current block corresponds to the projection location of the missing part' is not directly visualized or ablated, and the measured gain from the position-aware module is small, but this is an evidence/support weakness rather than circular reasoning; the claim is not derived from itself or from a fitted constant.
Assumptions & free parameters
free parameters (1)
- Per-block position weights (6 projection images x 2x2 blocks = 24 scalars) =
Not reported
assumptions (4)
- domain assumption Frozen CLIP embeddings trained on natural images transfer to textureless depth projection maps and provide useful geometric detail for point cloud completion.
- domain assumption EdgeConnect inpainting produces a reliable complete projection map for arbitrary partial depth projections.
- ad hoc to paper Textual templates of the form "There is {category} point cloud projection map" provide CLIP text features that improve geometric completion beyond the category already encoded in the point cloud.
- ad hoc to paper The learned 2x2 block weights correspond to the location of missing parts.
Cite this review
Pith. "Pith review of Position-aware Guided Point Cloud Completion with CLIP Model." pith.science (2026). https://pith.science/paper/6TIBTLS3
@misc{pith2026241208271,
author = {Pith},
title = {Pith review of: Position-aware Guided Point Cloud Completion with CLIP Model},
year = {2026},
howpublished = {\url{https://pith.science/paper/6TIBTLS3}},
note = {Machine review of arXiv:2412.08271}
}
read the original abstract
Point cloud completion aims to recover partial geometric and topological shapes caused by equipment defects or limited viewpoints. Current methods either solely rely on the 3D coordinates of the point cloud to complete it or incorporate additional images with well-calibrated intrinsic parameters to guide the geometric estimation of the missing parts. Although these methods have achieved excellent performance by directly predicting the location of complete points, the extracted features lack fine-grained information regarding the location of the missing area. To address this issue, we propose a rapid and efficient method to expand an unimodal framework into a multimodal framework. This approach incorporates a position-aware module designed to enhance the spatial information of the missing parts through a weighted map learning mechanism. In addition, we establish a Point-Text-Image triplet corpus PCI-TI and MVP-TI based on the existing unimodal point cloud completion dataset and use the pre-trained vision-language model CLIP to provide richer detail information for 3D shapes, thereby enhancing performance. Extensive quantitative and qualitative experiments demonstrate that our method outperforms state-of-the-art point cloud completion methods.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Aiello, E.; Valsesia, D.; and Magli, E. 2022. Cross-modal learning for image-guided point cloud shape completion. NeurIPS, 37349--37362
work page 2022
-
[4]
Cai, P.; Scott, D.; Li, X.; and Wang, S. 2024. Orthogonal Dictionary Guided Shape Completion Network for Point Cloud. In AAAI, 864--872
work page 2024
-
[5]
Chang, A. X.; Funkhouser, T.; Guibas, L.; Hanrahan, P.; Huang, Q.; Li, Z.; Savarese, S.; Savva, M.; Song, S.; Su, H.; et al. 2015. Shapenet: An information-rich 3d model repository. arXiv preprint arXiv:1512.03012
arXiv 2015
-
[6]
Chefer, H.; Gur, S.; and Wolf, L. 2021. Generic Attention-Model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers. In ICCV, 397--406
work page 2021
-
[7]
Chen, Y.-N.; Dai, H.; and Ding, Y. 2022. Pseudo-stereo for monocular 3d object detection in autonomous driving. In CVPR, 887--897
work page 2022
-
[8]
Chen, Z.; Long, F.; Qiu, Z.; Yao, T.; Zhou, W.; Luo, J.; and Mei, T. 2023. Anchorformer: Point cloud completion from discriminative nodes. In CVPR, 13581--13590
work page 2023
Show all 52 references
-
[9]
Cheng, X.; and Li, L. 2024. Open 3D World in Autonomous Driving. arXiv preprint arXiv:2408.10880
2024 arXiv
-
[10]
Christen, S.; Yang, W.; P \'e rez-D’Arpino, C.; Hilliges, O.; Fox, D.; and Chao, Y.-W. 2023. Learning human-to-robot handovers from point clouds. In CVPR, 9654--9664
2023
-
[11]
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. ICLR
2021
-
[12]
Geiger, A.; Lenz, P.; Stiller, C.; and Urtasun, R. 2013. Vision meets robotics: The kitti dataset. The International Journal of Robotics Research, 1231--1237
2013
-
[13]
Han, Z.; Chen, C.; Liu, Y.-S.; and Zwicker, M. 2020. ShapeCaptioner: Generative caption network for 3D shapes by learning a mapping from parts detected in multiple views to sentences. In ACM MM, 1018--1027
2020
-
[14]
Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.-T.; Parekh, Z.; Pham, H.; Le, Q.; Sung, Y.-H.; Li, Z.; and Duerig, T. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML, 4904--4916. PMLR
2021
-
[15]
Kasten, Y.; Rahamim, O.; and Chechik, G. 2024. Point cloud completion with pretrained text-to-image diffusion models. NeurIPS, 36
2024
-
[16]
Li, L. 2023. Hierarchical edge aware learning for 3d point cloud. In Computer Graphics International Conference, 81--92. Springer
2023
-
[17]
Li, S.; Gao, P.; Tan, X.; and Wei, M. 2023. Proxyformer: Proxy alignment assisted point cloud completion with missing part sensitive transformer. In CVPR, 9466--9475
2023
-
[18]
Mao, J.; Shi, S.; Wang, X.; and Li, H. 2023. 3D object detection for autonomous driving: A comprehensive survey. IJCV, 1909--1963
2023
-
[19]
Nazeri, K.; Ng, E.; Joseph, T.; Qureshi, F.; and Ebrahimi, M. 2019. Edgeconnect: Structure guided image inpainting using edge prediction. In ICCV
2019
-
[20]
T.; Hua, B.-S.; Tran, K.; Pham, Q.-H.; and Yeung, S.-K
Nguyen, D. T.; Hua, B.-S.; Tran, K.; Pham, Q.-H.; and Yeung, S.-K. 2016. A field model for repairing 3d shapes. In CVPR, 5676--5684
2016
-
[21]
C.; Nord-Larsen, T.; Trepekli, K.; Gieseke, F.; and Igel, C
Oehmcke, S.; Li, L.; Revenga, J. C.; Nord-Larsen, T.; Trepekli, K.; Gieseke, F.; and Igel, C. 2022. Deep learning based 3D point cloud regression for estimating forest biomass. In Proceedings of the 30th international conference on advances in geographic information systems, 1--4
2022
-
[22]
C.; Nord-Larsen, T.; Gieseke, F.; and Igel, C
Oehmcke, S.; Li, L.; Trepekli, K.; Revenga, J. C.; Nord-Larsen, T.; Gieseke, F.; and Igel, C. 2024. Deep point cloud regression for above-ground forest biomass estimation from airborne LiDAR. Remote Sensing of Environment, 302: 113968
2024
-
[23]
Pan, L. 2020. ECG: Edge-aware point cloud completion with graph convolution. IEEE Robotics and Automation Letters, 4392--4398
2020
-
[24]
Pan, L.; Chen, X.; Cai, Z.; Zhang, J.; Zhao, H.; Yi, S.; and Liu, Z. 2021. Variational relational point completion network. In CVPR, 8524--8533
2021
-
[25]
R.; Su, H.; Mo, K.; and Guibas, L
Qi, C. R.; Su, H.; Mo, K.; and Guibas, L. J. 2017 a . Pointnet: Deep learning on point sets for 3d classification and segmentation. In CVPR, 652--660
2017
-
[26]
R.; Yi, L.; Su, H.; and Guibas, L
Qi, C. R.; Yi, L.; Su, H.; and Guibas, L. J. 2017 b . Pointnet++: Deep hierarchical feature learning on point sets in a metric space. NeurIPS, 30
2017
-
[27]
W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021. Learning transferable visual models from natural language supervision. In ICML, 8748--8763
2021
-
[28]
Rong, Y.; Zhou, H.; Yuan, L.; Mei, C.; Wang, J.; and Lu, T. 2024. CRA-PCN: Point Cloud Completion with Intra-and Inter-level Cross-Resolution Transformers. In AAAI, 4676--4685
2024
-
[29]
Song, W.; Zhou, J.; Wang, M.; Tan, H.; Li, N.; and Liu, X. 2023. Fine-grained Text and Image Guided Point Cloud Completion with CLIP Model. arXiv preprint arXiv:2308.08754
2023 arXiv
-
[30]
Tang, J.; Gong, Z.; Yi, R.; Xie, Y.; and Ma, L. 2022. Lake-net: Topology-aware point cloud completion by localizing aligned keypoints. In CVPR, 1726--1735
2022
-
[31]
P.; Kosaraju, V.; Rezatofighi, H.; Reid, I.; and Savarese, S
Tchapmi, L. P.; Kosaraju, V.; Rezatofighi, H.; Reid, I.; and Savarese, S. 2019. Topnet: Structural point cloud decoder. In CVPR, 383--392
2019
-
[32]
Z.; and Yan, S
Wang, J.; Zhou, P.; Shou, M. Z.; and Yan, S. 2023. Position-guided text prompt for vision-language pre-training. In CVPR, 23242--23251
2023
-
[33]
H.; and Lee, G
Wang, X.; Ang Jr, M. H.; and Lee, G. H. 2020. Cascaded refinement network for point cloud completion. In CVPR, 790--799
2020
-
[34]
Wen, X.; Xiang, P.; Han, Z.; Cao, Y.-P.; Wan, P.; Zheng, W.; and Liu, Y.-S. 2021. Pmp-net: Point cloud completion by learning multi-step point moving paths. In CVPR, 7443--7452
2021
-
[35]
Wen, X.; Xiang, P.; Han, Z.; Cao, Y.-P.; Wan, P.; Zheng, W.; and Liu, Y.-S. 2022. Pmp-net++: Point cloud completion by transformer-enhanced multi-step point moving paths. TPAMI, 852--867
2022
-
[36]
Xiang, P.; Wen, X.; Liu, Y.-S.; Cao, Y.-P.; Wan, P.; Zheng, W.; and Han, Z. 2021. Snowflakenet: Point cloud completion by snowflake point deconvolution with skip-transformer. In ICCV, 5499--5509
2021
-
[37]
Xie, H.; Yao, H.; Zhou, S.; Mao, J.; Zhang, S.; and Sun, W. 2020. Grnet: Gridding residual network for dense point cloud completion. In ECCV, 365--381
2020
-
[38]
Yan, X.; Yan, H.; Wang, J.; Du, H.; Wu, Z.; Xie, D.; Pu, S.; and Lu, L. 2022. Fbnet: Feedback network for point cloud completion. In ECCV, 676--693
2022
-
[39]
Yang, Y.; Feng, C.; Shen, Y.; and Tian, D. 2018. Foldingnet: Point cloud auto-encoder via deep grid deformation. In CVPR, 206--215
2018
-
[40]
Yu, J.; Wang, Z.; Vasudevan, V.; Yeung, L.; Seyedhosseini, M.; and Wu, Y. 2022. CoCa: Contrastive Captioners are Image-Text Foundation Models. Trans. Mach. Learn. Res
2022
-
[41]
Yu, X.; Rao, Y.; Wang, Z.; Liu, Z.; Lu, J.; and Zhou, J. 2021. Pointr: Diverse point cloud completion with geometry-aware transformers. In ICCV, 12498--12507
2021
-
[42]
Yu, X.; Rao, Y.; Wang, Z.; Lu, J.; and Zhou, J. 2023. AdaPoinTr: Diverse Point Cloud Completion With Adaptive Geometry-Aware Transformers. TPAMI, 14114--14130
2023
-
[43]
Yuan, W.; Khot, T.; Held, D.; Mertz, C.; and Hebert, M. 2018. Pcn: Point completion network. In 3DV, 728--737
2018
-
[44]
Zhang, B.; Zhao, X.; Wang, H.; and Hu, R. 2022 a . Shape completion with points in the shadow. In SIGGRAPH Asia, 1--9
2022
-
[45]
Zhang, P.; Li, X.; Hu, X.; Yang, J.; Zhang, L.; Wang, L.; Choi, Y.; and Gao, J. 2021 a . VinVL: Making Visual Representations Matter in Vision-Language Models. CVPR
2021
-
[46]
Zhang, R.; Guo, Z.; Zhang, W.; Li, K.; Miao, X.; Cui, B.; Qiao, Y.; Gao, P.; and Li, H. 2022 b . Pointclip: Point cloud understanding by clip. In CVPR, 8552--8562
2022
-
[47]
Zhang, W.; Yan, Q.; and Xiao, C. 2020. Detail preserved point cloud completion via separated feature aggregation. In ECCV, 512--528
2020
-
[48]
Zhang, W.; Zhou, H.; Dong, Z.; Liu, J.; Yan, Q.; and Xiao, C. 2022 c . Point cloud completion via skeleton-detail transformer. TVCG, 29(10): 4229--4242
2022
-
[49]
Zhang, X.; Feng, Y.; Li, S.; Zou, C.; Wan, H.; Zhao, X.; Guo, Y.; and Gao, Y. 2021 b . View-guided point cloud completion. In CVPR, 15890--15899
2021
-
[50]
Zhou, H.; Cao, Y.; Chu, W.; Zhu, J.; Lu, T.; Tai, Y.; and Wang, C. 2022. Seedformer: Patch seeds based point cloud completion with upsample transformer. In ECCV, 416--432
2022
-
[51]
Zhu, Z.; Chen, H.; He, X.; Wang, W.; Qin, J.; and Wei, M. 2023 a . Svdformer: Complementing point cloud via self-view augmentation and self-structure dual-generator. In ICCV, 14508--14518
2023
-
[52]
Zhu, Z.; Nan, L.; Xie, H.; Chen, H.; Wang, J.; Wei, M.; and Qin, J. 2023 b . Csdn: Cross-modal shape-transfer dual-refinement network for point cloud completion. TVCG
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.