REVIEW 2 major objections 5 minor 66 references
Hierarchical Compact Clustering Attention (COCA) for Unsupervised Object-Centric Learning
T0 review · 2 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Compactness, not preset slots, drives this unsupervised object segmenter.
desk verdict COCA-Net is a genuinely new compactness-guided hierarchical clustering architecture for object-centric learning with strong synthetic results, but the abstract overclaims dynamic slots and the sequential clustering concealment is not what it says it is. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the compactness functional $C_i(\Lambda^l_i)$ of Eq. 3, which measures a candidate affinity mask's moment of inertia around node $i$ against that of a disk of equal area; because inertia is minimized at the true centroid, the node at an object's centroid receives the highest compactness score within that object. This single score drives anchor selection in the sequential stick-breaking clustering loop (Eqs. 5a-5d), which in turn determines all cluster assignments in the hierarchy.
What would settle it
Run COCA-Net unsupervised on a synthetic dataset of hollow rings, elongated rods, or U-shaped objects: if the highest-scoring anchor consistently lands away from the object's perceptual center, or on the background, the central premise fails and segmentation accuracy should drop sharply relative to the convex-blob datasets. The check can be made directly by comparing the argmax of Eq. 3 to the true object centroids on such shapes.
Extended reading notes
Core claim
COCA-Net claims that unsupervised object discovery reduces to a greedy search over compactness. A COCA layer refines pixel features with self-attention, computes an affinity mask for every node in a window, scores each mask with a mass-normalized moment-of-inertia compactness functional (Eq. 3), and then runs a stick-breaking concealing loop (Eqs. 5a-5d): take the highest-scoring remaining mask as the next cluster, mask out its nodes, and repeat until the scope is exhausted. Stacked in a bottom-up hierarchy over non-overlapping windows, this produces a dendrogram of object masks with a variable number of clusters, trained end-to-end with only a pixel-reconstruction loss. The paper reports state-of-the-art or competitive results on Tetrominoes, Multi-dSprites, ObjectsRoom, ShapeStacks, CLEVR6, and CLEVRTex across ARI and mean Segmentation Covering, on both the decoder side and, unusually, the encoder side, with background regions segmented more coherently than by the baselines.
Load-bearing premise
The method assumes that, within any window, the node whose affinity mask is most compact by the moment-of-inertia score is the centroid of a distinct foreground object, and that compactness separates objects from background; this premise is only validated on synthetic scenes of convex, blob-like objects.
Editorial extensions
If this is right
- Dynamic slot allocation: because clustering stops when the scope is exhausted or drops below a threshold, the number of output masks is not fixed in advance; the COCA-Net-Dyna variant trained on CLEVR6 transfers to CLEVR10 with only a small drop (ARI 0.978 vs 0.985).
- The encoder alone produces segmentation masks that match or beat decoder masks on most datasets, so the trained hierarchy could be reused as a standalone unsupervised object-centric feature extractor.
- Background elements are segmented as coherent clusters: with background included in the evaluation, COCA-Net gains roughly thirty ARI points over INV-SA and BOQ-SA on ObjectsRoom, where those baselines collapse to background ARI around 0.6 or lower.
- Training is more stable: over three seeds COCA-Net shows markedly smaller standard deviation than GEN-v2, INV-SA, and BOQ-SA on nearly every dataset and metric.
- The windowed hierarchy operates on $U\times U$ windows in parallel with per-layer complexity reducible to $O(N^2\log N)$, making the architecture amenable to parallel implementation and scaling to higher resolutions.
Reading between the lines
- A sharp empirical boundary follows from the compactness prior: the method should transfer to real-world images only insofar as objects are convex and blob-like; on elongated, hollow, or heavily concave objects the Eq. 3 anchor selection would likely drift from the perceptual center, a regime the six synthetic benchmarks do not probe.
- Because the dendrogram is produced deterministically from compactness scores, encoder-side masks could be emitted without a decoder forward pass, suggesting a cheap inference mode for downstream tasks that the paper hints at but does not test.
- The explicit geometric prior yields a falsifiable prediction that slot-attention models lack: deforming an object (stretching it, adding holes, making it concave) should degrade COCA-Net's anchor selection in a monotone, predictable way, and this degradation could be measured directly on synthetic shape morphs.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes the Compact Clustering Attention (COCA) layer, an attention-based clustering module for unsupervised object-centric learning. COCA scores candidate affinity masks by a moment-of-inertia compactness measure and sequentially selects the most compact mask as an object cluster using a stick-breaking concealment strategy. Stacking these layers forms COCA-Net, which is trained end-to-end with a spatial broadcast decoder. The authors evaluate on six synthetic datasets (Tetrominoes, Multi-dSprites, ObjectsRoom, ShapeStacks, CLEVR6, CLEVRTex) against GEN-v2, INVSA, and BOQSA, reporting decoder and encoder segmentation masks, with and without background, and include ablations for compactness-based versus random anchor selection and for dynamic slot allocation.
Significance. If the results hold, COCA-Net offers a credible alternative to slot-attention models, with particular strengths in encoder-side segmentation, background handling, and lower variance across training runs. The compactness functional is an external geometric measure from the cited literature, not fitted to the data, and evaluation is on held-out test images, so circularity concerns are limited. The ablations in Table 3 and supplementary Tables 9-10 provide machine-checked support for the central design choice of compactness-guided anchor selection. However, the conceptual novelty rests on the claimed stick-breaking concealment mechanism and on the compactness prior; both need closer scrutiny, as detailed below.
major comments (2)
- [3.1.5, Eqs. (5a)-(5d) and Algorithm 1] The compactness scores C^l are computed once per layer from the full affinity masks Λ^l (Eq. 3), before sequential clustering begins. Eq. (5a) only multiplies C_{m-1} by the scope Z_{m-1}; it does not recompute Λ^l or C^l on the residual set of unassigned nodes. Consequently, the sentence in Sec. 3.1.5 that this "ensuring that only unassigned nodes (with non-zero scope values) contribute to subsequent calculations" is inaccurate for the anchor-selection signal. A remaining node's compactness score still includes contributions from pixels already assigned to earlier clusters, so the selected anchor may not be the most compact residual shape. This issue is load-bearing because compactness is the only inductive bias for choosing object centroids. The authors should either modify the algorithm to recompute compactness after masking the affinity masks by the scope, or revise the text to state that scores are computed once per layer and provide an experiment or argument demonstrating that stale scores do not degrade anchor selection.
- [3.1.5, compactness assumption] The paper assumes, in Sec. 3.1.5, that nodes corresponding to the centroids of distinct objects yield the highest compactness scores within their objects, and that foreground objects are more compact than background regions. This assumption is validated only on synthetic datasets composed of convex, blob-like objects. For elongated, concave, or hollow objects, or when background regions are more compact than foreground, the anchor-selection mechanism would likely fail. Because this compactness prior is the central spatial inductive bias of the method, the authors should either restrict their claims to the considered benchmark, or include a stress test with non-convex shapes (e.g., synthetic objects with holes or elongated structures) to characterize the limits of the approach.
minor comments (5)
- [Abstract and Sec. 4.3] The abstract and the concluding paragraph of Sec. 4.3 claim that COCA-Net is "not bound by a predetermined number of object masks," but all main experiments fix the number of output slots to the dataset maximum, and only the COCA-Net-Dyna ablation supports the dynamic-slot claim. Please qualify this claim in the abstract and conclusion.
- [Supplementary, Algorithm 1, line 6] In Algorithm 1, line 6 of the supplementary material, the scope update is written as "Zl = Zl ⊙ (1− Πl−1)"; this appears to use an undefined Πl−1. It should be the current cluster mask Πm (or Π appended in the previous line). Please correct the notation.
- [Tables 3, 9, 10] Table 3 and the supplementary Tables 9-10 report single-seed results, whereas Tables 1 and 2 report mean ± standard deviation over three seeds. To avoid confusion when comparing these results, the captions should state explicitly that these are single-seed numbers.
- [Eq. (3)] Equation (3) includes a summation over j < v without defining the index v in the main text. Please define all summation indices and the range of v.
- [Sec. 3.1.4] The text states that "A perfect circle achieves the maximum compactness value of 1," but Eq. (3) evaluates compactness around an arbitrary node i; the maximum may be less than 1 when i is not the centroid of the shape. Please rephrase to clarify that the bound applies when the reference point is the centroid.
Circularity Check
No significant circularity: the compactness measure is an external geometric prior, and the segmentation claims are benchmarked against held-out ground truth.
full rationale
The paper's derivation chain is self-contained and non-circular. The compactness functional in Eq. 2 and Eq. 3 is an external geometric measure taken from references [39,54,55], not a fitted parameter and not defined in terms of the segmentation targets. The sequential clustering in Eqs. 5a-5d uses that external score to select anchors, and the output masks are evaluated against held-out ground truth on six synthetic datasets, so the central quantitative claims are external comparisons rather than re-statements of the method's inputs. No parameter is fitted to the reported ARI/mSC targets, and the ablation variants (COCA-Net-RAS and COCA-Net-Dyna) alter the algorithm rather than redefining the evaluation metric. The paper's compactness-as-objectness premise is an explicit inductive bias and a possible domain limitation for non-convex objects, but an inductive bias is not circularity: the predicted masks are not equal by construction to the compactness scores or to any fitted training target. The reviewer concern that masking does not recompute compactness on the residual set is a correctness or robustness question about the algorithm, not a circularity, because the output segmentation is still judged against external ground truth. There are no self-citations used as load-bearing evidence, and no uniqueness theorem or prior result by the same authors is invoked to force the method's choices.
Assumptions & free parameters
free parameters (6)
- Per-layer temperature tau_l =
0.75 to 2.00 across layers/datasets (Table 6)
- Number of COCA layers L =
2 or 3 depending on dataset (Table 6)
- Window sizes and per-layer output cluster counts (h_l,w_l,k_l) =
See Table 6, e.g., Tetrominoes [[4,4],[8,8]], k=[3,4]
- ViT-22B refinement stack depth and attention kernel sizes =
Table 6, e.g., 2 or 3 ViT layers per COCA layer depending on L
- Dyna stopping threshold =
2.5% of initial scope
- Pad and random crop augmentation amount =
three pixels
assumptions (4)
- standard math A node located at the centroid of its affinity mask's shape yields the maximum compactness score
- domain assumption Foreground objects in the evaluated scenes are more compact and convex than background regions
- domain assumption Objects can be recovered by clustering within non-overlapping windows at each hierarchy level and merging across levels
- domain assumption Soft-argmin with min-max scaling converts Euclidean feature distances into affinity masks whose compactness ranking is informative for objectness
Cite this review
Pith. "Pith review of Hierarchical Compact Clustering Attention (COCA) for Unsupervised Object-Centric Learning." pith.science (2026). https://pith.science/paper/TGO6ADEH
@misc{pith2026250502071,
author = {Pith},
title = {Pith review of: Hierarchical Compact Clustering Attention (COCA) for Unsupervised Object-Centric Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/TGO6ADEH}},
note = {Machine review of arXiv:2505.02071}
}
read the original abstract
We propose the Compact Clustering Attention (COCA) layer, an effective building block that introduces a hierarchical strategy for object-centric representation learning, while solving the unsupervised object discovery task on single images. COCA is an attention-based clustering module capable of extracting object-centric representations from multi-object scenes, when cascaded into a bottom-up hierarchical network architecture, referred to as COCA-Net. At its core, COCA utilizes a novel clustering algorithm that leverages the physical concept of compactness, to highlight distinct object centroids in a scene, providing a spatial inductive bias. Thanks to this strategy, COCA-Net generates high-quality segmentation masks on both the decoder side and, notably, the encoder side of its pipeline. Additionally, COCA-Net is not bound by a predetermined number of object masks that it generates and handles the segmentation of background elements better than its competitors. We demonstrate COCA-Net's segmentation performance on six widely adopted datasets, achieving superior or competitive results against the state-of-the-art models across nine different evaluation metrics.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Barron, Fer- ran Marques, and Jitendra Malik
Pablo Arbelaez, Jordi Pont-Tuset, Jonathan T. Barron, Fer- ran Marques, and Jitendra Malik. Multiscale combinatorial grouping. In Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), 2014. 3
work page 2014
-
[2]
Contour detection and hierarchical image seg- mentation
Pablo Arbel ´aez, Michael Maire, Charless Fowlkes, and Ji- tendra Malik. Contour detection and hierarchical image seg- mentation. IEEE Transactions on Pattern Analysis and Ma- chine Intelligence, 33(5):898–916, 2011. 3
work page 2011
-
[3]
CuVLER: Enhanced Unsupervised Object Discov- eries through Exhaustive Self-Supervised Transformers
Shahaf Arica, Or Rubin, Sapir Gershov, and Shlomi Laufer. CuVLER: Enhanced Unsupervised Object Discov- eries through Exhaustive Self-Supervised Transformers . In 2024 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 23105–23114, Los Alami- tos, CA, USA, 2024. IEEE Computer Society. 3
work page 2024
-
[4]
Spectral clustering with graph neural networks for graph pooling
Filippo Maria Bianchi, Daniele Grattarola, and Cesare Alippi. Spectral clustering with graph neural networks for graph pooling. In International Conference on Machine Learning, 2019. 3
work page 2019
-
[5]
Ondrej Biza, Sjoerd Van Steenkiste, Mehdi S. M. Saj- jadi, Gamaleldin Fathy Elsayed, Aravindh Mahendran, and Thomas Kipf. Invariant slot attention: Object discovery with slot-centric reference frames. In Proceedings of the 40th In- ternational Conference on Machine Learning , pages 2507–
-
[6]
Burgess, Lo ¨ıc Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matthew M
Christopher P. Burgess, Lo ¨ıc Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matthew M. Botvinick, and Alexander Lerchner. Monet: Unsupervised scene decompo- sition and representation. ArXiv, abs/1901.11390, 2019. 3, 5
arXiv 1901
-
[7]
HASSOD: Hierarchical adaptive self-supervised ob- ject detection
Shengcao Cao, Dhiraj Joshi, Liangyan Gui, and Yu-Xiong Wang. HASSOD: Hierarchical adaptive self-supervised ob- ject detection. In Thirty-seventh Conference on Neural In- formation Processing Systems, 2023. 3
work page 2023
-
[8]
SOHES: Self-supervised open- world hierarchical entity segmentation
Shengcao Cao, Jiuxiang Gu, Jason Kuen, Hao Tan, Ruiyi Zhang, Handong Zhao, Ani Nenkova, Liangyan Gui, Tong Sun, and Yu-Xiong Wang. SOHES: Self-supervised open- world hierarchical entity segmentation. In The Twelfth In- ternational Conference on Learning Representations , 2024. 3
work page 2024
Show all 66 references
-
[9]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv’e J’egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9630–9640, 2021. 3
2021
-
[10]
Spotlight atten- tion: Robust object-centric learning with a spatial locality prior
Ayush K Chakravarthy, Trang Nguyen, Anirudh Goyal, Yoshua Bengio, and Michael Curtis Mozer. Spotlight atten- tion: Robust object-centric learning with a spatial locality prior. ArXiv, abs/2305.19550, 2023. 2
2023 arXiv
-
[11]
Object representations as fixed points: Training iterative refinement algorithms with implicit differentiation
Michael Chang, Tom Griffiths, and Sergey Levine. Object representations as fixed points: Training iterative refinement algorithms with implicit differentiation. In Advances in Neu- ral Information Processing Systems , pages 32694–32708. Curran Associates, Inc., 2022. 1, 2
2022
-
[12]
Spatially invariant unsuper- vised object detection with convolutional neural networks
Eric Crawford and Joelle Pineau. Spatially invariant unsuper- vised object detection with convolutional neural networks. In AAAI Conference on Artificial Intelligence, 2019. 3
2019
-
[13]
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Peter Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdul- mohsin, Rodolphe Jenatton, Lucas Beyer, Michael Tschan- nen, Anurag Arnab, Xiao Wang, Carlos Riquelme Ruiz, M...
2023
-
[14]
Perceptual group to- kenizer: Building perception with iterative grouping
Zhiwei Deng, Ting Chen, and Yang Li. Perceptual group to- kenizer: Building perception with iterative grouping. In The Twelfth International Conference on Learning Representa- tions, 2024. 3
2024
-
[15]
Dittadi, S
A. Dittadi, S. S. Papa, M. De Vita, B. Sch¨olkopf, O. Winther, and F. Locatello. Generalization and robustness implications in object-centric learning. In Proceedings of the 39th In- ternational Conference on Machine Learning (ICML), pages 5221–5285. PMLR, 2022. 1, 2, 6, 16
2022
-
[16]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...
2021
-
[17]
Kosiorek, Oiwi Parker Jones, and Ingmar Posner
Martin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, and Ingmar Posner. Genesis: Generative scene inference and sampling with object-centric latent representations. ArXiv, abs/1907.13052, 2019. 3, 6
1907 arXiv
-
[18]
Genesis-v2: Inferring unordered object representations with- out iterative refinement
Martin Engelcke, Oiwi Parker Jones, and Ingmar Posner. Genesis-v2: Inferring unordered object representations with- out iterative refinement. In Neural Information Processing Systems, 2021. 3, 5, 6, 7, 8, 16, 18
2021
-
[19]
Everitt, S
B. Everitt, S. Landau, M. Leese, D. Stahl, and an O’Reilly Media Company Safari. Cluster Analysis, 5th Edition. John Wiley & Sons, 2011. 2
2011
-
[20]
Adap- tive slot attention: Object discovery with dynamic slot num- ber, 2024
Ke Fan, Zechen Bai, Tianjun Xiao, Tong He, Max Horn, Yanwei Fu, Francesco Locatello, and Zheng Zhang. Adap- tive slot attention: Object discovery with dynamic slot num- ber, 2024. 2
2024
-
[21]
Gonzalez and Richard E
Rafael C. Gonzalez and Richard E. Woods. Digital Image Processing (3rd Edition). Prentice-Hall, Inc., USA, 2006. 2
2006
-
[22]
On the binding problem in artificial neural networks
Klaus Greff, Sjoerd van Steenkiste, and J¨urgen Schmidhuber. On the binding problem in artificial neural networks. ArXiv, abs/2012.05208, 2020. 1, 2, 3
2012 arXiv
-
[23]
Fuchs, Ingmar Posner, and Andrea Vedaldi
Oliver Groth, Fabian B. Fuchs, Ingmar Posner, and Andrea Vedaldi. Shapestacks: Learning vision-based physical intu- ition for generalised object stacking. In Computer Vision – ECCV 2018, pages 724–739, Cham, 2018. Springer Interna- tional Publishing. 6 9
2018
-
[24]
Vision GNN: An image is worth graph of nodes
Kai Han, Yunhe Wang, Jianyuan Guo, Yehui Tang, and En- hua Wu. Vision GNN: An image is worth graph of nodes. In Advances in Neural Information Processing Systems , 2022. 3
2022
-
[25]
Hastie, R
T. Hastie, R. Tibshirani, and J.H. Friedman. The Elements of Statistical Learning: Data Mining, Inference, and Predic- tion. Springer, 2009. 2
2009
-
[26]
Comparing partitions
Lawrence Hubert and Phipps Arabie. Comparing partitions. Journal of Classification, 2(1):193–218, 1985. 6
1985
-
[27]
Jain and Richard C
Anil K. Jain and Richard C. Dubes. Algorithms for clustering data. 1988. 2
1988
-
[28]
Improving object- centric learning with query optimization
Baoxiong Jia, Yu Liu, and Siyuan Huang. Improving object- centric learning with query optimization. InThe Eleventh In- ternational Conference on Learning Representations , 2023. 1, 2, 6, 7, 8, 16, 18
2023
-
[29]
Scalor: Generative world models with scal- able object representations
Jindong Jiang, Sepehr Janghorbani, Gerard de Melo, and Sungjin Ahn. Scalor: Generative world models with scal- able object representations. In International Conference on Learning Representations, 2019. 3
2019
-
[30]
Lawrence Zitnick, and Ross Girshick
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elemen- tary visual reasoning. In 2017 IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR),...
2017
-
[31]
Multi-object datasets
Rishabh Kabra, Chris Burgess, Loic Matthey, Raphael Lopez Kaufman, Klaus Greff, Malcolm Reynolds, and Alexander Lerchner. Multi-object datasets. https://github.com/deepmind/multi-object-datasets/, 2019. 6
2019
-
[32]
The Background Also Matters: Background- Aware Motion-Guided Objects Discovery
Sandra Kara, Hejer Ammar, Florian Chabot, and Quoc- Cuong Pham. The Background Also Matters: Background- Aware Motion-Guided Objects Discovery . In 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 1205–1214, Los Alamitos, CA, USA,
2024
-
[33]
Kaufman and P.J
L. Kaufman and P.J. Rousseeuw. Finding Groups in Data: An Introduction to Cluster Analysis. Wiley, 2009. 2
2009
-
[34]
Tsung-Wei Ke, Jyh-Jing Hwang, Yunhui Guo, Xudong Wang, and Stella X. Yu. Unsupervised hierarchical semantic segmentation with multiview cosegmentation and clustering transformers. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 2561–2571,
2022
-
[35]
Tsung-Wei Ke, Sangwoo Mo, and Stella X. Yu. Learning hi- erarchical image segmentation for recognition and by recog- nition. In The Twelfth International Conference on Learning Representations, 2024. 3
2024
-
[36]
Shepherding slots to objects: Towards stable and ro- bust object-centric learning
Jinwoo Kim, Janghyuk Choi, Ho-Jin Choi, and Seon Joo Kim. Shepherding slots to objects: Towards stable and ro- bust object-centric learning. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19198–19207, 2023. 1, 2
2023
-
[37]
Space: Unsupervised object-oriented scene representa- tion via spatial attention and decomposition
Zhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun, Gautam Singh, Fei Deng, Jindong Jiang, and Sungjin Ahn. Space: Unsupervised object-oriented scene representa- tion via spatial attention and decomposition. InInternational Conference on Learning Representations, 2020. 3
2020
-
[38]
Object- centric learning with slot attention
Francesco Locatello, Dirk Weissenborn, Thomas Un- terthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf. Object- centric learning with slot attention. In Advances in Neural Information Processing Systems, pages 11525–11538. Cur- ran...
2020
-
[39]
Massam and Michael F
Bryan H. Massam and Michael F. Goodchild. Temporal trends in the spatial organization of a service agency. Cana- dian Geographies / G ´eographies canadiennes , 15(3):193– 206, 1971. 5
1971
-
[40]
Oswald, Cees G
Duy-Kien Nguyen, Mahmoud Assran, Unnat Jain, Martin R. Oswald, Cees G. M. Snoek, and Xinlei Chen. An image is worth more than 16x16 patches: Exploring transformers on individual pixels. ArXiv, abs/2406.09415, 2024. 3
2024 arXiv
-
[41]
Delving into the whorl of flower segmentation
Maria-Elena Nilsback and Andrew Zisserman. Delving into the whorl of flower segmentation. In British Machine Vision Conference, 2007. 17
2007
-
[42]
Comprehensive survey on hierarchical clus- tering algorithms and the recent developments
Xingcheng Ran, Yue Xi, Yonggang Lu, Xiangwen Wang, and Zhenyu Lu. Comprehensive survey on hierarchical clus- tering algorithms and the recent developments. Artificial In- telligence Review, 56:8219 – 8264, 2022. 2, 6
2022
-
[43]
Bridging the gap to real-world object-centric learning
Maximilian Seitzer, Max Horn, Andrii Zadaianchuk, Do- minik Zietlow, Tianjun Xiao, Carl-Johann Simon-Gabriel, Tong He, Zheng Zhang, Bernhard Sch¨olkopf, Thomas Brox, and Francesco Locatello. Bridging the gap to real-world object-centric learning. In The Eleventh International ...
2023
-
[44]
Normalized cuts and image segmentation
Jianbo Shi and Jitendra Malik. Normalized cuts and image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 22, 2002. 3
2002
-
[45]
Illiterate DALL- e learns to compose
Gautam Singh, Fei Deng, and Sungjin Ahn. Illiterate DALL- e learns to compose. In International Conference on Learn- ing Representations, 2022. 17
2022
-
[46]
Graph clustering with graph neural net- works
Anton Tsitsulin, John Palowitch, Bryan Perozzi, and Em- manuel M ¨uller. Graph clustering with graph neural net- works. J. Mach. Learn. Res., 24(1), 2024. 3
2024
-
[47]
Selective search for object recognition
J R R Uijlings, K E A van de Sande, T Gevers, and A W M Smeulders. Selective search for object recognition. Interna- tional Journal of Computer Vision , 104(2):154–171, 2013. 3
2013
-
[48]
Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Neural Infor- mation Processing Systems, 2017. 3, 13
2017
-
[49]
Yu, and Ishan Misra
Xudong Wang, Rohit Girdhar, Stella X. Yu, and Ishan Misra. Cut and learn for unsupervised object detection and instance segmentation. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 3124–3134,
2023
-
[50]
Seg- ment anything without supervision
Xudong Wang, Jingfeng Yang, and Trevor Darrell. Seg- ment anything without supervision. In The Thirty-eighth An- nual Conference on Neural Information Processing Systems,
-
[51]
Crowley, and Dominique Vaufrey- daz
Yangtao Wang, Xi Shen, Yuan Yuan, Yuming Du, Maomao Li, Shell Xu Hu, James L. Crowley, and Dominique Vaufrey- daz. Tokencut: Segmenting objects in images and videos with self-supervised transformer and normalized cut. IEEE 10 Trans. Pattern Anal. Mach. Intell. , 45(12):15790–15801,
-
[52]
Burgess, and Alexander Lerchner
Nicholas Watters, Lo ¨ıc Matthey, Christopher P. Burgess, and Alexander Lerchner. Spatial broadcast decoder: A simple ar- chitecture for learning disentangled representations in vaes. ArXiv, abs/1901.07017, 2019. 2
1901 arXiv
-
[53]
Belongie, and Pietro Perona
Peter Welinder, Steve Branson, Takeshi Mita, Catherine Wah, Florian Schroff, Serge J. Belongie, and Pietro Perona. Caltech-ucsd birds 200. 2010. 17
2010
-
[54]
Wentz Wenwen Li, Tingyong Chen and Chao Fan
Elizabeth A. Wentz Wenwen Li, Tingyong Chen and Chao Fan. Nmmi: A mass compactness measure for spatial pat- tern analysis of areal features. Annals of the Association of American Geographers, 104(6):1116–1133, 2014. 5
2014
-
[55]
Goodchild Wenwen Li and Richard Church
Michael F. Goodchild Wenwen Li and Richard Church. An efficient measure of compactness for two-dimensional shapes and its application in regionalization problems. In- ternational Journal of Geographical Information Science, 27 (6):1227–1250, 2013. 5, 14
2013
-
[56]
Group normalization
Yuxin Wu and Kaiming He. Group normalization. In Pro- ceedings of the European Conference on Computer Vision (ECCV), 2018. 4
2018
-
[57]
Rui Xu and Donald C. Wunsch. Survey of clustering al- gorithms. IEEE Transactions on Neural Networks, 16:645– 678, 2005. 2
2005
-
[58]
Benchmarking and analysis of unsupervised object segmentation from Real-World single images
Yafei Yang and Bo Yang. Benchmarking and analysis of unsupervised object segmentation from Real-World single images. International Journal of Computer Vision , 132(6): 2077–2113, 2024. 1
2024
-
[59]
Zhang, Simon Lacoste-Julien, Gert- jan J
Yan Zhang, David W. Zhang, Simon Lacoste-Julien, Gert- jan J. Burghouts, and Cees G. M. Snoek. Unlocking slot at- tention by changing optimal transport costs. In Proceedings of the 40th International Conference on Machine Learning , pages 41931–41951. PMLR, 2023. 2
2023
-
[60]
Zimmermann, Sjoerd van Steenkiste, Mehdi S
Roland S. Zimmermann, Sjoerd van Steenkiste, Mehdi S. M. Sajjadi, Thomas Kipf, and Klaus Greff. Sensitivity of slot- based object-centric models to their number of slots. CoRR, abs/2305.18890, 2023. 2 11 Hierarchical Compact Clustering Attention (COCA) for Unsupervised Object-...
2023 arXiv
-
[63]
At each layer, this feature image is partitioned intoU×U non-overlapping windows
Complexity Analysis Suppose that COCA-Net is to process a feature image con- sisting ofN×N pixels. At each layer, this feature image is partitioned intoU×U non-overlapping windows. The com- plexity burden of a COCA layer resides on the generation of affinity masks and their co...
-
[64]
all-to-all
Architectural Details 7.1. Pixel Feature Encoder We use a simple backbone to initialize pixel features that encode both appearance and positional information. Specif- ically, the input image is first processed through a sin- gle convolutional layer that preserves resolution to...
-
[65]
pad and random crop
Implementation Details For all experiments included in this work, we build on the comprehensive OCL library that is provided by [15]. This OCL library includes the six datasets that we share results on, in addition to code scripts for training, testing and eval- uation. Detail...
2000
-
[66]
Additional Results 9.1. Initial Experiments on Real-World Datasets; Birds and Flowers We evaluated COCA-Net on Birds [53] and Flowers [41] (following BOQSA’s protocol) using the SLATE [45] ar- chitecture which leverages a transformer decoder but no pretrained backbone. Entire ...
-
[2024]
IEEE Computer Society. 2
-
[2527]
2, 6, 7, 8, 16, 18
PMLR, 2023. 2, 6, 7, 8, 16, 18
2023
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.