REVIEW 3 major objections 5 minor 48 references
Unsupervised Part Discovery via Descriptor-Based Masked Image Restoration with Optimized Constraints
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read MPAE discovers meaningful object parts with no labels: filling randomly masked patches with learned part descriptors and restoring the image aligns descriptor semantics with real part shapes, and the method outperforms baselines on…
desk verdict Genuinely new mechanism for unsupervised part discovery, but the paper's own ablation shows the reported NMI/ARI metrics reward a degenerate foreground-segmentation solution, so the quantitative claims need a credibility fix. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Masked Part Autoencoder and its matching block. A frozen self-supervised Vision Transformer plus a trainable $1 \times 1$ convolution produces a dense feature map $F$; the model learns $K+1$ part descriptors $D$ (one per foreground part plus background); softmax similarity $P_{i,j,k}$ scores each pixel against each descriptor; masked positions of the filled feature map $R$ are set to $\sum_{k=1}^{K+1} P_{i,j,k} D_k$, while visible positions keep the unmasked patch features $F^U$. Restoration with an L1 term plus a perceptual term then pulls descriptors and visible features into shared latent clusters. Three constraints stabilize this: a presence constraint (each part appears somewhere in a mini-group, background near image borders), a semantic consistency constraint with additive angular margin on descriptor-to-region cosine similarity, and a distribution constraint combining total variation and entropy to sharpen part boundaries.
What would settle it
Train MPAE with a decoder strong enough to reconstruct masked patches almost perfectly from visible context alone, for example by using a much deeper decoder or a very low masking ratio, and then measure whether the similarity maps still match annotated part boundaries. If reconstruction stays accurate while the similarity maps no longer track part shapes, the claimed implicit alignment through restoration is not the driving mechanism.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that masked image restoration can function as a part-alignment mechanism: by using learned part descriptors to fill masked patches and then restoring the image, the model implicitly clusters the filled descriptor features together with unmasked patch features inside each part region, while low-level appearance from visible patches guides the descriptors to the correct part boundaries. The resulting similarity maps can be read directly as pixel-level part masks. The paper further argues that looser constraints, which allow parts to be absent from individual images, let the same descriptors transfer across categories and handle occlusion, and it supports this with experiments showing improved normalized mutual information and adjusted Rand index over existing methods on four benchmarks.
Load-bearing premise
The load-bearing premise is that filling masked patches with part descriptors and restoring the image will, by itself, pull descriptors and unmasked patch features into shared latent clusters whose similarity maps track part shapes; the paper's ablation shows this alignment is fragile, since removing the perceptual term drops normalized mutual information (NMI) from 55.10 to 19.65 on the multi-category segmentation benchmark.
Editorial extensions
If this is right
- With no labels, MPAE can produce part-level masks that follow real object boundaries, which is a prerequisite for part-level features in downstream tasks such as fine-grained recognition and person search.
- Because parts may be absent from any single image, the same training procedure transfers across categories, allowing the model to share similar parts such as wheels, heads, or hulls across different object classes.
- The method is designed to handle occlusion, since it learns which parts are present rather than assuming all parts appear in every image.
- The optimal masking ratio around 90 percent indicates that the alignment mechanism depends on forcing most of the reconstruction through part descriptors rather than through visible context alone.
Reading between the lines
- A testable extension follows from the masking-ratio curve: if descriptor alignment is driven by restoration pressure, then an annealing schedule that raises the masking ratio during training could improve convergence and final part quality, though the paper does not test this.
- The role of the perceptual loss suggests the method's success is tied to structural supervision in pixel space; one could test whether a latent-space perceptual loss, or a stronger structural criterion, changes how tightly similarity maps follow boundaries.
- The mini-group presence constraint is an implicit batch-level prior; the paper's ablation shows very large groups fragment rare regions into spurious parts, so a principled choice of group size could be connected to part frequency statistics across categories.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes MPAE, a masked autoencoder variant for unsupervised part discovery. The model learns K+1 part descriptors from a trainable ViT, extracts dense features from a frozen pretrained ViT, and uses the similarity between descriptors and local features to fill masked patches in an image restoration objective. Several auxiliary losses are introduced: a presence loss (Lf and Lb), a semantic consistency loss (Ls), and a distribution loss (Ld). The paper reports NMI/ARI comparisons on PartImageNet, CUB, and CelebA, plus extensive ablations over masking ratios, hyperparameters, backbones, and loss components. The code is provided.
Significance. The idea of using masked restoration with part descriptors as fill tokens is interesting, and the qualitative results suggest plausible part masks across categories and scenarios. The paper ships code and includes extensive ablations, which is a strength. If the quantitative claims held, this would be a useful step toward unsupervised part discovery. However, the evaluation issues detailed below mean that the numerical superiority over prior work is not currently established, and the paper's own ablations indicate that the primary metric may reward a degenerate foreground/background solution.
major comments (3)
- [Section 4.4, Table 4] The paper's primary evidence is NMI/ARI, yet its own ablation shows these metrics reward a degenerate solution. Removing Ls increases NMI/ARI at K=8 (55.69/76.54 vs 33.51/65.05) and K=25 (57.70/76.97 vs 37.22/68.04), while the text states the model without Ls 'degrades into a foreground segmentation model' and Appendix B.8 reports only 3.94 foreground parts per image. Thus the higher NMI/ARI of the full model at K=50 (55.10/73.52) could in principle be driven by coarse foreground/background separation rather than meaningful part structure. The comparisons in Tables 1-3 are therefore contaminated: they do not establish that MPAE discovers parts better than a simple foreground segmentation model. Please add baselines that explicitly evaluate part-level quality (e.g., part-aware IoU or semantic part consistency) and report the number of active parts per image for all methods.
- [Table 1 and Table 6] The comparison with Xia et al. is at a different cluster count (K=4Nc, i.e., 436 for OOD and 632 for Segmentation) versus MPAE at K=8/25/50. NMI and ARI are known to depend on the number of clusters, so these numbers are not directly comparable. Notably, Table 6 shows that changing MPAE's K from 50 to 4Nc reduces NMI from 55.10 to 47.59 and ARI from 73.52 to 70.71, which is a large drop. Please either evaluate MPAE at the same K as Xia et al., or evaluate Xia et al. at the same K as MPAE, and discuss the cluster-count dependence.
- [Section 3.1.1 and Appendix B.1] The paper asserts that the restoration objective 'implicitly clusters' the filled descriptors and unmasked patch features, and that this is the mechanism that aligns descriptors with part shapes. The only direct evidence is the large drop when the VGG-19 structural penalty is removed (NMI 55.10 to 19.65, ARI 73.52 to 49.72 in Table 7). This shows the component matters but does not show that the clustering actually occurs as described. Please provide direct evidence, such as an analysis of the latent space distribution of filled versus unmasked features, or a controlled experiment that varies restoration difficulty, or temper the mechanistic claim to what is demonstrated.
minor comments (5)
- [Section 1 and Section 4.1] Typos: 'inpus' should be 'inputs'; 'datasests' should be 'datasets'.
- [Contributions list] The phrase 'as followings' should be 'as follows'.
- [Equation (7)] The typesetting of the ArcFace-style loss is garbled (e.g., 'M kes' appears to be M_k * exp(...)). Please rewrite with clear mathematical notation and define all symbols in one place.
- [Section 3.3] The term 'PartFormer' is used without a formal definition; earlier in Section 3.1.1 it is mentioned in passing. Please introduce the term when the architecture is first described.
- [Section 4.4, masking ratio paragraph] The sentence 'a higher r results in more features in the filled feature map R comes from part descriptors D' has a subject-verb agreement problem and should be rephrased.
Circularity Check
No circularity: MPAE's central claims are tested against external benchmarks with standard training losses; the cited self-work is design inheritance, not a load-bearing derivation.
full rationale
The paper's derivation chain is not circular at the equation level. The training objective, Eq. (3), is a masked image restoration loss combining L1 reconstruction with a VGG-19 perceptual penalty, and the auxiliary constraints in Eqs. (4)-(11) are standard presence, semantic-consistency, total-variation, and entropy losses. None of these losses contains the evaluation metrics NMI, ARI, or NME; the reported numbers are computed against external human annotations on PartImageNet, CUB, and CelebA. The part descriptors D and similarity maps P are learned parameters and data-dependent outputs, not fitted values that are later renamed as predictions. The self-citations to the authors' prior TPAMI work [43] are limited to architectural choices and hyperparameters: 'as [43], we replace the class embedding of a classical ViT with K + 1 embeddings' and 'Referring to [43], we set s and m to 20 and 0.5.' These borrowings are implementation details, and the paper actively compares against [43] and reports better numbers, so the conclusion does not rest on an unverified self-citation. The ablation in Table 4, where removing Ls improves NMI/ARI while the paper states the model 'degrades into a foreground segmentation model,' is a measurement-validity concern rather than circularity: it says the metric can reward coarse foreground splits, but it does not show that the full model's result is equivalent by construction to its inputs. The mechanism claim in Section 3.1.1 that restoration 'implicitly clusters FU and D within the same part region together in latent space' is an empirical assumption supported by ablations, not a definitional equivalence. No uniqueness theorem, fitted-input-as-prediction, or ansatz-smuggled-via-citation pattern is present.
Assumptions & free parameters
free parameters (7)
- K (number of parts) =
8, 25, 50 (PartImageNet); 4, 8, 16 (CUB, CelebA)
- Masking ratio r =
0.90
- Mini-group size G =
64 (multi-category), 8 (single-category)
- Loss weights lambda_d, lambda_p, lambda_s =
0.5, 1.0, 0.25
- ArcFace scale s and margin m =
20 and 0.5
- Presence threshold for M_k =
0.001
- MPAE encoder/decoder layer counts =
2 layers each
assumptions (4)
- domain assumption The softmax similarity map P between learned descriptors and pretrained features can serve as a pixel-level part mask after training.
- domain assumption Masked restoration implicitly clusters descriptor-filled features with unmasked features of the same part in latent space.
- domain assumption Each mini-group of G images contains all K foreground parts and background appears near image boundaries.
- domain assumption NMI/ARI between predicted clusters and annotated part masks measures part discovery quality.
Cite this review
Pith. "Pith review of Unsupervised Part Discovery via Descriptor-Based Masked Image Restoration with Optimized Constraints." pith.science (2026). https://pith.science/paper/EJ3ULHWT
@misc{pith2026250711985,
author = {Pith},
title = {Pith review of: Unsupervised Part Discovery via Descriptor-Based Masked Image Restoration with Optimized Constraints},
year = {2026},
howpublished = {\url{https://pith.science/paper/EJ3ULHWT}},
note = {Machine review of arXiv:2507.11985}
}
read the original abstract
Part-level features are crucial for image understanding, but few studies focus on them because of the lack of fine-grained labels. Although unsupervised part discovery can eliminate the reliance on labels, most of them cannot maintain robustness across various categories and scenarios, which restricts their application range. To overcome this limitation, we present a more effective paradigm for unsupervised part discovery, named Masked Part Autoencoder (MPAE). It first learns part descriptors as well as a feature map from the inputs and produces patch features from a masked version of the original images. Then, the masked regions are filled with the learned part descriptors based on the similarity between the local features and descriptors. By restoring these masked patches using the part descriptors, they become better aligned with their part shapes, guided by appearance features from unmasked patches. Finally, MPAE robustly discovers meaningful parts that closely match the actual object shapes, even in complex scenarios. Moreover, several looser yet more effective constraints are proposed to enable MPAE to identify the presence of parts across various scenarios and categories in an unsupervised manner. This provides the foundation for addressing challenges posed by occlusion and for exploring part similarity across multiple categories. Extensive experiments demonstrate that our method robustly discovers meaningful parts across various categories and scenarios. The code is available at the project https://github.com/Jiahao-UTS/MPAE.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
Deep vit features as dense visual descriptors
Shir Amir, Yossi Gandelsman, Shai Bagon, and Tali Dekel. Deep vit features as dense visual descriptors. InECCV Work- shops, 2022. 1, 3, 6, 7, 13
work page 2022
-
[2]
Pdiscoformer: Relaxing part discovery constraints with vision transformers
Ananthu Aniraj, Cassio F Dantas, Dino Ienco, and Diego Marcos. Pdiscoformer: Relaxing part discovery constraints with vision transformers. In ECCV, 2024. 1, 2, 4, 6, 7, 8, 11
work page 2024
-
[3]
Unsupervised learn- ing of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Pi- otr Bojanowski, and Armand Joulin. Unsupervised learn- ing of visual features by contrasting cluster assignments. In NIPS, pages 9912–9924. Curran Associates, Inc., 2020. 2
work page 2020
-
[4]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e Jegou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In ICCV, pages 9630–9640, 2021. 1, 2, 13
work page 2021
-
[5]
An empiri- cal study of training self-supervised vision transformers
Xinlei Chen, Saining Xie, and Kaiming He. An empiri- cal study of training self-supervised vision transformers. In ICCV, pages 9620–9629, 2021. 2
work page 2021
-
[6]
Unsupervised part discovery from con- trastive reconstruction
Subhabrata Choudhury, Iro Laina, Christian Rupprecht, and Andrea Vedaldi. Unsupervised part discovery from con- trastive reconstruction. In NIPS, pages 28104–28118, 2021. 1, 3, 7
work page 2021
-
[7]
Deep feature factorization for concept discovery
Edo Collins, Radhakrishna Achanta, and Sabine S ¨usstrunk. Deep feature factorization for concept discovery. In ECCV, pages 352–368, 2018. 3
work page 2018
-
[8]
Vision transformers need registers
Timoth ´ee Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski. Vision transformers need registers. In ICLR.,
Show all 48 references
-
[9]
Arcface: Additive angular mar- gin loss for deep face recognition
Jiankang Deng, Jia Guo, Jing Yang, Niannan Xue, Irene Kot- sia, and Stefanos Zafeiriou. Arcface: Additive angular mar- gin loss for deep face recognition. IEEE TPAMI, 44(10): 5962–5979, 2022. 5
2022
-
[10]
BERT: Pre-training of deep bidirectional trans- formers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional trans- formers for language understanding. InNAACL, pages 4171– 4186, 2019. 2
2019
-
[11]
Adam: A method for stochastic optimization
Kingma Diederik and Ba Jimmy. Adam: A method for stochastic optimization. In ICLR, 2015. 11
2015
-
[12]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...
2021
-
[13]
Un- supervised co-part segmentation through assembly
Qingzhe Gao, Bin Wang, Libin Liu, and Baoquan Chen. Un- supervised co-part segmentation through assembly. InICML, pages 3576–3586, 2021. 1, 2, 3
2021
-
[14]
Bootstrap your own latent - a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altch ´e, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Ghesh- laghi Azar, Bilal Piot, koray kavukcuoglu, Remi Munos, and Michal Valko. Bootstrap your own latent - a ne...
-
[15]
Siamese masked autoencoders
Agrim Gupta, Jiajun Wu, Jia Deng, and Fei-Fei Li. Siamese masked autoencoders. In NIPS, pages 40676–40693. Curran Associates, Inc., 2023. 2
2023
-
[16]
Partimagenet: A large, high- quality dataset of parts
Ju He, Shuo Yang, Shaokang Yang, Adam Kortylewski, Xi- aoding Yuan, Jie-Neng Chen, Shuai Liu, Cheng Yang, Qi- hang Yu, and Alan Yuille. Partimagenet: A large, high- quality dataset of parts. In ECCV, pages 128–145, 2022. 5, 11
2022
-
[18]
Momentum contrast for unsupervised visual rep- resentation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual rep- resentation learning. In CVPR, pages 9726–9735, 2020. 2
2020
-
[19]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In CVPR, pages 15979–15988, 2022. 2, 3, 4
2022
-
[20]
Ganseg: Learning to segment by unsupervised hierarchical image generation
Xingzhe He, Bastian Wandt, and Helge Rhodin. Ganseg: Learning to segment by unsupervised hierarchical image generation. In CVPR, pages 1215–1225, 2022. 1, 2, 7
2022
-
[21]
Interpretable and accurate fine- grained recognition via region grouping
Zixuan Huang and Yin Li. Interpretable and accurate fine- grained recognition via region grouping. In CVPR, pages 8659–8669, 2020. 2, 6, 7, 11
2020
-
[22]
Scops: Self-supervised co-part segmentation
Wei-Chih Hung, Varun Jampani, Sifei Liu, Pavlo Molchanov, Ming-Hsuan Yang, and Jan Kautz. Scops: Self-supervised co-part segmentation. In CVPR, pages 869–878, 2019. 1, 3, 5, 8, 11
2019
-
[23]
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In ECCV, pages 694–711, 2016. 4
2016
-
[24]
Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C. Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick. Segment anything. In ICCV, pages 3992– 4003, 2023. 14
2023
-
[25]
Unsupervised part segmentation through disentangling ap- pearance and shape
Shilong Liu, Lei Zhang, Xiao Yang, Hang Su, and Jun Zhu. Unsupervised part segmentation through disentangling ap- pearance and shape. In CVPR, pages 8351–8360, 2021. 1, 2, 3
2021
-
[26]
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In ICCV, pages 3730–3738, 2015. 5, 11
2015
-
[27]
Animal kingdom: A large and diverse dataset for animal behavior understanding
Xun Long Ng, Kian Eng Ong, Qichen Zheng, Yun Ni, Si Yong Yeo, and Jun Liu. Animal kingdom: A large and diverse dataset for animal behavior understanding. InCVPR, pages 19001–19012, 2022. 1
2022
-
[28]
Maxime Oquab, Timoth ´ee Darcet, Th´eo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael ...
2024
-
[29]
Plot: Text-based person search with part slot attention for corresponding part discovery
Jicheol Park, Dongwon Kim, Boseung Jeong, and Suha Kwak. Plot: Text-based person search with part slot attention for corresponding part discovery. In ECCV, pages 474–490,
-
[30]
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. OpenAI blog, 2018. 2
2018
-
[31]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In ICML, pages 8748–8763, 2021. 14
2021
-
[32]
Rudin, Stanley Osher, and Emad Fatemi
Leonid I. Rudin, Stanley Osher, and Emad Fatemi. Nonlinear total variation based noise removal algorithms. Physica D: Nonlinear Phenomena, 60(1):259–268, 1992. 5
1992
-
[33]
Particle: Part discov- ery and contrastive learning for fine-grained recognition
Oindrila Saha and Subhransu Maji. Particle: Part discov- ery and contrastive learning for fine-grained recognition. In ICCV Workshops, pages 167–176, 2023. 2
2023
-
[34]
Jianbo Shi and J. Malik. Normalized cuts and image seg- mentation. IEEE TPAMI, 22(8):888–905, 2000. 2
2000
-
[35]
Motion- supervised co-part segmentation
Aliaksandr Siarohin, Subhankar Roy, St ´ephane Lathuili `ere, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. Motion- supervised co-part segmentation. In ICPR, pages 9650– 9657, 2021. 1, 3
2021
-
[36]
Very deep convo- lutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. Very deep convo- lutional networks for large-scale image recognition. InICLR,
-
[37]
Unsu- pervised learning of object landmarks by factorized spatial embeddings
James Thewlis, Hakan Bilen, and Andrea Vedaldi. Unsu- pervised learning of object landmarks by factorized spatial embeddings. In ICCV, pages 3229–3238, 2017. 3
2017
-
[38]
van der Klis, S
R. van der Klis, S. Alaniz, M. Mancini, C. F. Dantas, D. Ienco, Z. Akata, and D. Marcos. Pdisconet: Semantically consistent part discovery for fine-grained recognition. In ICCV, pages 1866–1876, 2023. 1, 2, 6, 7, 11
2023
-
[39]
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. The Caltech-UCSD Birds-200-2011 Dataset. Technical Re- port CNS-TR-2011-001, California Institute of Technology,
2011
-
[40]
Yu, and Ishan Misra
Xudong Wang, Rohit Girdhar, Stella X. Yu, and Ishan Misra. Cut and learn for unsupervised object detection and instance segmentation. In CVPR, pages 3124–3134, 2023. 2
2023
-
[41]
Videocutler: Surprisingly simple unsuper- vised video instance segmentation
Xudong Wang, Ishan Misra, Ziyun Zeng, Rohit Girdhar, and Trevor Darrell. Videocutler: Surprisingly simple unsuper- vised video instance segmentation. In CVPR, pages 22755– 22764, 2024. 2
2024
-
[42]
Crowley, and Dominique Vaufreydaz
Yangtao Wang, Xi Shen, Shell Xu Hu, Yuan Yuan, James L. Crowley, and Dominique Vaufreydaz. Self-supervised trans- formers for unsupervised object discovery using normalized cut. In CVPR, pages 14523–14533, 2022. 2
2022
-
[43]
Unsupervised part dis- covery via dual representation alignment
Jiahao Xia, Wenjian Huang, Min Xu, Jianguo Zhang, Haimin Zhang, Ziyu Sheng, and Dong Xu. Unsupervised part dis- covery via dual representation alignment. IEEE TPAMI, pages 1–18, 2024. 1, 2, 3, 6, 7, 8, 11, 13
2024
-
[44]
Hp-capsule: Unsupervised face part discovery by hierarchical parsing capsule network
Chang Yu, Xiangyu Zhu, Xiaomei Zhang, Zidu Wang, Zhaoxiang Zhang, and Zhen Lei. Hp-capsule: Unsupervised face part discovery by hierarchical parsing capsule network. In CVPR, pages 4022–4031, 2022. 1, 2, 3
2022
-
[45]
How mask matters: Towards theoretical understandings of masked autoencoders
Qi Zhang, Yifei Wang, and Yisen Wang. How mask matters: Towards theoretical understandings of masked autoencoders. In NIPS, pages 27127–27139, 2022. 2
2022
-
[46]
Unsupervised discovery of object land- marks as structural representations
Yuting Zhang, Yijie Guo, Yixin Jin, Yijun Luo, Zhiyuan He, and Honglak Lee. Unsupervised discovery of object land- marks as structural representations. In CVPR, pages 2694– 2703, 2018. 1, 3
2018
-
[47]
Pass: Part-aware self-supervised pre- training for person re-identification
Kuan Zhu, Haiyun Guo, Tianyi Yan, Yousong Zhu, Jinqiao Wang, and Ming Tang. Pass: Part-aware self-supervised pre- training for person re-identification. In ECCV, pages 198– 214, 2022. 2
2022
-
[48]
Adrian Ziegler and Yuki M. Asano. Self-supervised learning of object parts for semantic segmentation. In ICCV, pages 14482–14491, Los Alamitos, CA, USA, 2022. 2 10 Appendix A.1 Dataset Details PartImageNet OOD dataset. We use the OOD variant of PartImageNet [16] to validate MP...
2022
-
[49]
Performance comparison of MPAE with different supervised pretrained backbones on the PartImageNet Segmentation dataset in the setting of K = 50
with [8] SAM [24] CLIP [31] NMI (%) ↑ 55.10 17.16 33.16 ARI (%) ↑ 73.32 55.26 72.90 Table 12. Performance comparison of MPAE with different supervised pretrained backbones on the PartImageNet Segmentation dataset in the setting of K = 50. The number of MPAE encoder and decoder...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.