REVIEW 3 major objections 5 minor 76 references
Symbolic Disentangled Representations for Images
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read ArSyD stores each image property as a full-size hypervector so editing an object is a vector swap.
desk verdict ArSyD is a promising HDC-based disentanglement architecture with real novelty, but the DCM metric is misimplemented and the 'by construction' claim overreaches. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is attention over a frozen random codebook (item memory), used as a bridge between the encoder's localist feature vector and a distributed hypervector representation. For each generative factor $i$, a projection of the encoder output is matched by softmax attention against fixed seed hypervectors from the codebook, and the weighted sum of those seeds becomes the factor's value vector $V_i^*$. These value vectors are bound to factor hypervectors and bundled into the object representation, following Holographic Reduced Representation operations. The same encoder runs on donor and target images, and the feature-exchange module swaps one value vector before decoding.
What would settle it
On the dSprites dataset, take two images that differ only in orientation, swap only the orientation value vector in the latent representation, decode both edited images, and classify their shape with a pretrained shape classifier. If the shape prediction changes when only orientation is exchanged, the orientation vector carries shape information and the factor-level disentanglement claim fails; the paper's own Figure 7 already shows such a shape distortion for orientation.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that a disentangled representation can be built as a sum of learned value hypervectors, one per generative factor, instead of a vector whose individual coordinates each encode a factor. Each value vector $V_i^*$ is produced by an attention mechanism that selects a weighted combination of fixed random seed vectors from an item memory, and the object representation is $O = \sum_i G_i \circledast V_i^*$. The model is trained by swapping the value vector of one factor between two images that differ only in that factor and reconstructing both images with an MSE loss. The authors argue that because each generative factor has its own vector and the value vectors are grounded in the image through attention, disentanglement holds by construction, and editing reduces to exchanging the corresponding vector.
Load-bearing premise
The paper's claim depends on the assumption that training only the attention weights, with nothing but reconstruction error on image pairs that differ in one factor, forces each selected hypervector to align with exactly one true generative factor and to stay independent of the others.
Editorial extensions
If this is right
- Editing an image property becomes a vector substitution, so the same edit operation is well-defined regardless of where in the latent space a factor lives.
- The model can be combined with Slot Attention to edit a single object inside a multi-object scene, not just isolated objects.
- Because disentanglement is claimed by construction, the approach avoids distributional assumptions and the loss-tuning typical of beta-VAE-style methods.
- The proposed DMM and DCM metrics allow comparisons between localist and distributed representations by measuring changes in pixel-space classifications.
- Reconstruction from a single factor vector shows that individual value vectors carry property-specific information, which supports interpretable editing.
Reading between the lines
- Editorial inference: if the value hypervectors are truly independent symbols, then latent vector arithmetic (for instance, adding the vector for one object property and subtracting another) should produce coherent composite edits, a test the paper does not run.
- Editorial inference: the entanglement of orientation with shape visible in the paper's Figure 7 suggests that adding a sparsity or orthogonality penalty on the attention weights would make the factor vectors cleaner; the paper itself does not propose such a penalty.
- Editorial inference: because DMM and DCM measure disentanglement through classifiers on reconstructed images, better decoders or stronger classifiers could change the measured scores even if the latent representation is unchanged, so the metrics should be read as bounded by the reconstruction and classification pipeline.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes ArSyD, an architecture that learns 'symbolic disentangled representations' in which each generative factor is represented by a full hypervector, and the object representation is a superposition of factor-value vectors obtained by attention over a fixed random codebook. The model is trained with weak supervision: image pairs differ in exactly one factor, and the feature-exchange reconstruction loss (Eq. 5) is the only training signal. The paper also introduces two classifier-based metrics, DMM and DCM, intended to evaluate disentanglement and compactness for distributed latent representations. Experiments on dSprites, CLEVR1, CLEVR5 (with Slot Attention), and a CelebA proof-of-concept show qualitative controlled edits and quantitative comparisons against BetaVAE and FactorVAE on FID/IoU, DMM, and DCM.
Significance. If the central claims held, this would be a valuable contribution: it offers a genuinely different, VSA-based route to disentangled representations that supports interpretable vector-level editing and proposes model-agnostic metrics that do not assume localist coordinates. The paper is also carefully structured around explicit research questions, and the qualitative exchange results on dSprites and CLEVR are compelling proof-of-concept. However, the evaluation instrument DCM is implemented inconsistently with its definition, and the 'disentanglement by construction' claim is stronger than the training objective and the paper's own qualitative results support. The significance of the contribution therefore depends on the extent to which these issues can be resolved.
major comments (3)
- [Section 3.4, Eq. (8), Figure 4] Eq. (8) defines DCM as (1/|C|) Σ_{c∈C} | Σ_{Gi∈G} [Cl(Ŝ'_c) ≠ ŷ]_{Gi} − 1 |, which is a per-unit row sum over generative factors. However, the text and Figure 4 state that DCM should use a per-factor column sum, counting how many latent units affect each generative factor. The metric's stated definition, 'whether each generative factor is encoded by a single latent unit,' requires a column sum. As written, DCM measures the average number of factors changed by a unit, not compactness across units for a factor. Consequently, the DCM values in Tables 3 and 4, including the conclusions that ArSyD is more compact than BetaVAE and FactorVAE, do not support the claims. Please correct Eq. (8) to aggregate over units for each factor, recompute the affected tables, and state clearly which operationalization is used.
- [Abstract; Section 3.1, Eqs. (3)–(4); Section 3.2–3.3, Eq. (5); Section 4, Figure 7] The abstract and Section 3.1 claim that 'disentanglement is achieved by construction,' but this is not established by the training procedure. In each weakly supervised pair, all factors except the exchanged one are identical between target and donor, so the exchange loss in Eq. (5) only requires V*_p to carry enough information to transfer factor p; it never penalizes V*_p for also encoding other factors, because those factors are constant in the pair and the decoder can ignore redundant information. The HDC binding in Eq. (1) makes the bound terms quasi-orthogonal, but it does not constrain V*_i itself to be independent of the other factors. The paper's own Figure 7 and accompanying text say that the 'Orientation' feature is 'strongly related' to the 'Shape' feature, which is direct evidence that the factor-alignment assumption can fail. The claim should be weakened to 'encouraged by the weak supervision objective,' or the authors should provide a direct analysis (e.g., probing or decoding each V*_i in isolation, or a full pairwise intervention matrix) demonstrating that each V*_i is factor-pure.
- [Section 4, RQ8, Figures 13–14] RQ8 concludes that the value vectors V*_i represent separate properties because decoding a single vector bound to a placeholder does not reconstruct a complete image. This conclusion does not follow from the evidence: a vector that encodes a mixture of several factors, but is alone insufficient to generate a full image, would produce qualitatively the same incomplete reconstructions. To support the factor-purity interpretation, the authors should test interventions directly, e.g., by modifying V*_i and measuring which downstream factor classifications change, or by training linear probes on V*_i to see whether they predict only the intended factor.
minor comments (5)
- [Figure 4] The caption says 'DCN metric' in the last sentence; this should be 'DCM metric.'
- [Eq. (5)] The text after Eq. (5) says '˜Od – a reconstructed target object'; this should be '˜Ot – a reconstructed target object,' and the donor terms should be labeled consistently.
- [Section 3.7 and Table 3] Section 3.7 states that dSprites paired was trained for 600 epochs, whereas the caption of Table 3 says 200 epochs for dSprites paired; please reconcile these numbers.
- [Section 3.8] The metric classifiers are described as six ResNet-34 models fine-tuned on the CLEVR1 paired dataset, but the paper also reports DMM and DCM for dSprites paired; please specify which classifiers are used for dSprites and CelebA, and report their per-factor accuracies on the reconstruction sets.
- [Section 3.4] The term 'unit' is introduced in the metrics section, but Figure 4 and the surrounding text use 'DCN' and 'DCM' inconsistently; please unify the notation.
Circularity Check
The headline 'disentanglement by construction' is partly definitional: the decomposition is named a 'symbolic disentangled representation' without an equation or loss term forcing each V*_i to be factor-pure; the empirical evaluations themselves are not fitted to the metrics.
-
self definitional
[Abstract; Section 3.1, paragraph after Eq. (4) and text defining the representation]
"Disentanglement is achieved by construction, no additional assumptions about the underlying distributions are made during training... The resulting value vectors V ∗ i are multiplied by the corresponding generative factor vectors Gi and summed to produce the vector O, which is a symbolic disentangled representation of an object."
The 'achieved by construction' claim reads the disentanglement property off the notation 'generative factor vector' and the name 'symbolic disentangled representation'. Equation (4) only defines V*_i as an attention mixture over the codebook of factor i; Equation (5) only requires MSE reconstruction after exchanging one index. Nothing in Eqs. (3)-(5) constrains V*_i to exclude information about other generative factors, and the paper itself reports in Figure 7 that on dSprites 'the Orientation feature is strongly related to the Shape feature'. Thus the asserted guarantee is not derived from the architecture or loss; it is equivalent to the paper's own definitional label for a sum of bound hypervectors.
full rationale
I examined the claimed derivation chain: HDC composition (Eq. 1-2), grounding through attention over a frozen codebook (Eq. 3-4), weak-supervision exchange with reconstruction-only loss (Eq. 5-6), and the DMM/DCM evaluation (Eq. 7-8). The only load-bearing step that reduces to its own input is the abstract's 'disentanglement is achieved by construction': the representation is called a 'symbolic disentangled representation' because it is written as a sum of vectors named after generative factors, but the learned V*_i are not proven factor-pure, and the paper's own qualitative results show orientation/shape entanglement. That is a terminological/definitional circularity in the central claim. The remaining machinery is not circular in the same way: DMM and DCM are not part of the training losses, the classifier-based metrics use external ResNet classifiers and are not fitted to ArSyD, and the comparisons also include IoU/FID against BetaVAE and FactorVAE. Self-citations in the paper ([21,22,29-32]) are contextual and not load-bearing. The Limitations section candidly acknowledges classifier dependence and the need to know generative factors beforehand; those are validity caveats, not circularity. Overall the central guarantee is overstated by definitional naming, but the empirical core retains independent content, so a moderate score is appropriate.
Assumptions & free parameters
free parameters (5)
- Latent dimension D =
1024 in main runs; swept from 16 to 2048
- Number of generative factors N =
5 for dSprites, 6 for CLEVR, 7 for CelebA
- Value codebook sizes per factor =
dSprites: 3 shapes, 6 scales, 40 orientations, 32 x-values, 32 y-values; CLEVR: per-dataset factor value lists…
- Attention K/Q projection size =
1024
- Number of slots K for CLEVR5 =
6
assumptions (6)
- standard math Random Gaussian hypervectors of dimension D are mutually quasi-orthogonal and circular convolution binding and unbinding behave as in HRR.
- domain assumption Each image contains one object with a known finite set of generative factors, or in CLEVR5 a known maximum number of objects and reliable slot decomposition.
- domain assumption Training pairs differ in exactly one known generative factor and the feature exchange vector e is available.
- domain assumption MSE reconstruction is a sufficient training signal for the attention outputs to become factor-aligned.
- domain assumption Classifiers trained on original images give reliable factor predictions on reconstructed images for DMM and DCM.
- domain assumption Slot Attention discovers object masks that are accurate enough for edited objects to be re-inserted into scenes.
invented entities (2)
-
Symbolic disentangled representation (superposition of bound hypervectors as a latent image code)
-
DMM and DCM disentanglement metrics
Cite this review
Pith. "Pith review of Symbolic Disentangled Representations for Images." pith.science (2026). https://pith.science/paper/TPIJX6E7
@misc{pith2026241219847,
author = {Pith},
title = {Pith review of: Symbolic Disentangled Representations for Images},
year = {2026},
howpublished = {\url{https://pith.science/paper/TPIJX6E7}},
note = {Machine review of arXiv:2412.19847}
}
read the original abstract
The idea of disentangled representations is to reduce the data to a set of generative factors that produce it. Typically, such representations are vectors in latent space, where each coordinate corresponds to one of the generative factors. The object can then be modified by changing the value of a particular coordinate, but it is necessary to determine which coordinate corresponds to the desired generative factor -- a difficult task if the vector representation has a high dimension. In this article, we propose ArSyD (Architecture for Symbolic Disentanglement), which represents each generative factor as a vector of the same dimension as the resulting representation. In ArSyD, the object representation is obtained as a superposition of the generative factor vector representations. We call such a representation a \textit{symbolic disentangled representation}. We use the principles of Hyperdimensional Computing (also known as Vector Symbolic Architectures), where symbols are represented as hypervectors, allowing vector operations on them. Disentanglement is achieved by construction, no additional assumptions about the underlying distributions are made during training, and the model is only trained to reconstruct images in a weakly supervised manner. We study ArSyD on the dSprites and CLEVR datasets and provide a comprehensive analysis of the learned symbolic disentangled representations. We also propose new disentanglement metrics that allow comparison of methods using latent representations of different dimensions. ArSyD allows to edit the object properties in a controlled and interpretable way, and the dimensionality of the object property representation coincides with the dimensionality of the object representation itself.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
Yoshua Bengio, Aaron Courville, and Pascal Vincent. Representation learning: A review and new perspectives.IEEE Transactions on Pattern Analysis and Machine Intelligence , 35(8):1798–1828, 2013. doi: 10. 1109/TPAMI.2013.50
work page 2013
-
[2]
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In F. Pereira, C.J. Burges, L. Bottou, and K.Q. Weinberger, editors, Advances in Neural Information Processing Systems , volume 25. Curran Asso- ciates, Inc., 2012. URL https://proceedings.neurips.cc/paper/2012/file/ c399862d3b9d6b76c84...
work page 2012
-
[3]
Efficient estimation of word representations in vector space
Tomás Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. In Yoshua Bengio and Yann LeCun, editors, ICLR, Workshop Track, 2013
work page 2013
-
[4]
Deep convolutional net- works on graph-structured data
Mikael Henaff, Joan Bruna, and Yann LeCun. Deep convolutional net- works on graph-structured data. ArXiv, abs/1506.05163, 2015
arXiv 2015
-
[5]
Petar Veli ˇckovi´c, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. Graph attention networks. In ICLR, 2018
work page 2018
-
[6]
beta-V AE: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. beta-V AE: Learning basic visual concepts with a constrained variational framework. In International Conference on Learning Representations,
-
[7]
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel. Infogan: Interpretable representation learning by information maximizing generative adversarial nets. In D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, editors, NIPS. Curran Associates, Inc., 2016
work page 2016
-
[8]
Disentangling by factorising
Hyunjik Kim and Andriy Mnih. Disentangling by factorising. In ICML, 2018
2018
Show all 76 references
-
[9]
Cian Eastwood and Christopher K. I. Williams. A framework for the quantitative evaluation of disentangled representations. In 6th Inter- national Conference on Learning Representations, ICLR 2018, Van- couver, BC, Canada, April 30 - May 3, 2018, Conference Track Pro- ceedings....
2018
-
[10]
Tell, draw, and repeat: Generating and modifying images based on continual linguistic instruction
Alaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, Devon Hjelm, Layla El Asri, Samira Ebrahimi Kahou, Yoshua Bengio, and Graham W.Taylor. Tell, draw, and repeat: Generating and modifying images based on continual linguistic instruction. ICCV, 2019
2019
-
[11]
Unified questioner trans- former for descriptive question generation in goal-oriented visual dia- logue
Shoya Matsumori, Kosuke Shingyouchi, Yukikoko Abe, Yosuke Fukuchi, Komei Sugiura, and Michita Imai. Unified questioner trans- former for descriptive question generation in goal-oriented visual dia- logue. ICCV, 2021
2021
-
[12]
Fast exploration with simplified models and approximately optimistic plan- ning in model based reinforcement learning, 2018
Ramtin Keramati, Jay Whang, Patrick Cho, and Emma Brunskill. Fast exploration with simplified models and approximately optimistic plan- ning in model based reinforcement learning, 2018
2018
-
[13]
Unsu- pervised learning of object keypoints for perception and control
Tejas D Kulkarni, Ankush Gupta, Catalin Ionescu, Sebastian Borgeaud, Malcolm Reynolds, Andrew Zisserman, and V olodymyr Mnih. Unsu- pervised learning of object keypoints for perception and control. Ad- vances in neural information processing systems, 32, 2019
2019
-
[14]
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemyslaw Debiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Christopher Hesse, Rafal Józefowicz, Scott Gray, Catherine Olsson, Jakub W. Pachocki, Michael Petrov, Henrique Pondé de Oliveira Pinto...
1912 arXiv
-
[15]
Object-centric task and motion planning in dynamic environments.IEEE Robotics and Automation Let- ters, 5:844–851, 2020
Toki Migimatsu and Jeannette Bohg. Object-centric task and motion planning in dynamic environments.IEEE Robotics and Automation Let- ters, 5:844–851, 2020
2020
-
[16]
Are disentangled representations helpful for abstract visual reasoning? In NeurIPS, 2019
Sjoerd van Steenkiste, Francesco Locatello, Jürgen Schmidhuber, and Olivier Bachem. Are disentangled representations helpful for abstract visual reasoning? In NeurIPS, 2019
2019
-
[17]
Systematic visual reasoning through object-centric relational abstraction
Taylor Whittington Webb, Shanka Subhra Mondal, and Jonathan Co- hen. Systematic visual reasoning through object-centric relational abstraction. In Thirty-seventh Conference on Neural Information Processing Systems , 2023. URL https://openreview.net/forum?id= 8JCZe7QrPy
2023
-
[18]
Shapestacks: Learning vision-based physical intuition for generalised object stacking
Oliver Groth, Fabian B Fuchs, Ingmar Posner, and Andrea Vedaldi. Shapestacks: Learning vision-based physical intuition for generalised object stacking. In Proceedings of the european conference on com- puter vision (eccv), pages 702–717, 2018
2018
-
[19]
Tenenbaum
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B. Tenenbaum. Clevrer: Collision events for video representation and reasoning. ArXiv, abs/1910.01442, 2020
1910 arXiv
-
[20]
Illiterate dall-e learns to compose
Gautam Singh, Fei Deng, and Sungjin Ahn. Illiterate dall-e learns to compose. ArXiv, abs/2110.11405, 2021
2021 arXiv
-
[21]
Quantized disentangled representa- tions for object-centric visual tasks
Daniil Kirilenko, Alexandr Korchemnyi, Konstantin Smirnov, Alexey K Kovalev, and Aleksandr I Panov. Quantized disentangled representa- tions for object-centric visual tasks. InInternational Conference on Pat- tern Recognition and Machine Intelligence , pages 514–522. Springer, 2023
2023
-
[22]
Object-centric learning with slot mixture module
Daniil Kirilenko, Vitaliy V orobyov, Alexey Kovalev, and Aleksandr Panov. Object-centric learning with slot mixture module. InThe Twelfth International Conference on Learning Representations , 2024. URL https://openreview.net/forum?id=aBUidW4Nkd
2024
-
[23]
Hyperdimensional computing: An introduction to com- puting in distributed representation with high-dimensional random vec- tors
Pentti Kanerva. Hyperdimensional computing: An introduction to com- puting in distributed representation with high-dimensional random vec- tors. Cognitive Computation , 1(2):139–159, Jun 2009. ISSN 1866-
2009
-
[24]
Rachkovskij, Evgeny Osipov, and Abbas Rahimi
Denis Kleyko, Dmitri A. Rachkovskij, Evgeny Osipov, and Abbas Rahimi. A survey on hyperdimensional computing aka vector symbolic architectures, part i: Models and data transformations. ACM Comput. Surv., 55(6), December 2022. ISSN 0360-0300. doi: 10.1145/3538531. URL https://d...
2022 doi
-
[25]
A survey on hyperdimensional computing aka vector symbolic ar- chitectures, part ii: Applications, cognitive models, and challenges
Denis Kleyko, Dmitri Rachkovskij, Evgeny Osipov, and Abbas Rahimi. A survey on hyperdimensional computing aka vector symbolic ar- chitectures, part ii: Applications, cognitive models, and challenges. ACM Comput. Surv. , 55(9), January 2023. ISSN 0360-0300. doi: 10.1145/3558000...
2023 doi
-
[26]
Concepts as semantic pointers: A framework and computational model
Peter Blouw, Eugene Solodkin, Paul Thagard, and Chris Eliasmith. Concepts as semantic pointers: A framework and computational model. Cognitive science , 40 5:1128–62, 2016. URL https://api. semanticscholar.org/CorpusID:16809232
2016
-
[27]
Distributed representation of n-gram statistics for boosting self-organizing maps with hyperdi- mensional computing
Denis Kleyko, Evgeny Osipov, Daswin De Silva, Urban Wiklund, Va- leriy Vyatkin, and Damminda Alahakoon. Distributed representation of n-gram statistics for boosting self-organizing maps with hyperdi- mensional computing. In Nikolaj Bjørner, Irina Virbitskaite, and An- drei V o...
2019
-
[28]
Vector symbolic architec- tures for context-free grammars
Peter beim Graben, Markus Huber, Werner Meyer, Ronald Römer, Constanze Tschöpe, and Matthias Wolff. Vector symbolic architec- tures for context-free grammars. CoRR, abs/2003.05171, 2020. URL https://arxiv.org/abs/2003.05171
2003 arXiv
-
[29]
Applying vector symbolic architecture and semi- otic approach to visual dialog
Alexey K Kovalev, Makhmud Shaban, Anfisa A Chuganskaya, and Aleksandr I Panov. Applying vector symbolic architecture and semi- otic approach to visual dialog. In Hybrid Artificial Intelligent Systems: 16th International Conference, HAIS 2021, Bilbao, Spain, September 22–24, 20...
2021
-
[30]
Question answering for visual navigation in human-centered environments
Daniil E Kirilenko, Alexey K Kovalev, Evgeny Osipov, and Aleksandr I Panov. Question answering for visual navigation in human-centered environments. In Mexican International Conference on Artificial Intel- ligence, pages 31–45. Springer, 2021
2021
-
[31]
Kovalev, Makhmud Shaban, Evgeny Osipov, and Alek- sandr I
Alexey K. Kovalev, Makhmud Shaban, Evgeny Osipov, and Alek- sandr I. Panov. Vector semiotic model for visual question answer- ing. Cognitive Systems Research, 71:52–63, 2022. ISSN 1389-0417. doi: https://doi.org/10.1016/j.cogsys.2021.09.001. URL https://www. sciencedirect.com/...
2022 doi
-
[32]
Vector symbolic scene representation for semantic place recognition
Daniil Kirilenko, Alexey K Kovalev, Yaroslav Solomentsev, Alexander Melekhin, Dmitry A Yudin, and Aleksandr I Panov. Vector symbolic scene representation for semantic place recognition. In 2022 Interna- tional Joint Conference on Neural Networks (IJCNN), pages 1–8. IEEE, 2022
2022
-
[33]
Analogy making and logical inference on images using cellular automata based hyperdimensional computing, 2015
Ozgur Yilmaz. Analogy making and logical inference on images using cellular automata based hyperdimensional computing, 2015
2015
-
[34]
Hyperseed: Unsupervised learning with vector symbolic ar- chitectures
Evgeny Osipov, Sachin Kahawala, Dilantha Haputhanthri, Thimal Kempitiya, Daswin De Silva, Damminda Alahakoon, and Denis Kleyko. Hyperseed: Unsupervised learning with vector symbolic ar- chitectures. CoRR, abs/2110.08343, 2021. URL https://arxiv.org/abs/ 2110.08343
-
[35]
The concentration of measure phenomenon
Michel Ledoux. The concentration of measure phenomenon. AMS Sur- veys and Monographs, 89, 01 2001
2001
-
[36]
Neural ma- chine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural ma- chine translation by jointly learning to align and translate. In Yoshua Bengio and Yann LeCun, editors, 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Confer...
2015 arXiv
-
[37]
Weakly supervised disentanglement with guarantees
Rui Shu, Yining Chen, Abhishek Kumar, Stefano Ermon, and Ben Poole. Weakly supervised disentanglement with guarantees. arXiv preprint arXiv:1910.09772, 2019
1910 arXiv
-
[38]
Weakly-supervised disentan- glement without compromises
Francesco Locatello, Ben Poole, Gunnar Rätsch, Bernhard Schölkopf, Olivier Bachem, and Michael Tschannen. Weakly-supervised disentan- glement without compromises. InInternational Conference on Machine Learning, pages 6348–6359. PMLR, 2020
2020
-
[39]
Object-centric learning with slot attention
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf. Object-centric learning with slot attention. Advances in neural information processing systems, 33:11525–11538, 2020
2020
-
[40]
dsprites: Disentanglement testing sprites dataset
Loic Matthey, Irina Higgins, Demis Hassabis, and Alexan- der Lerchner. dsprites: Disentanglement testing sprites dataset. https://github.com/deepmind/dsprites-dataset/, 2017
2017
-
[41]
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In CVPR, 2017
2017
-
[42]
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of International Conference on Computer Vision (ICCV), December 2015
2015
-
[43]
Burgess, Irina Higgins, Arka Pal, Loic Matthey, Nick Watters, Guillaume Desjardins, and Alexander Lerchner
Christopher P. Burgess, Irina Higgins, Arka Pal, Loic Matthey, Nick Watters, Guillaume Desjardins, and Alexander Lerchner. Understand- ing disentangling in beta-vae, 2018. URL https://arxiv.org/abs/1804. 03599
2018
-
[44]
Ricky T. Q. Chen, Xuechen Li, Roger B Grosse, and David K Duvenaud. Isolating sources of disentanglement in variational au- toencoders. In S. Bengio, H. Wallach, H. Larochelle, K. Grau- man, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neu- ral Information Processing ...
2018
-
[45]
Vari- ational inference of disentangled latent concepts from unlabeled obser- vations
Abhishek Kumar, Prasanna Sattigeri, and Avinash Balakrishnan. Vari- ational inference of disentangled latent concepts from unlabeled obser- vations. In 6th International Conference on Learning Representations, ICLR. OpenReview.net, 2018
2018
-
[46]
Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Gen- erative adversarial networks, 2014. URL https://arxiv.org/abs/1406. 2661
2014
-
[47]
Infogan-cr: Disentangling generative adversarial networks with contrastive regularizers
Zinan Lin, Kiran Koshy Thekumparampil, Giulia Fanti, and Sewoong Oh. Infogan-cr: Disentangling generative adversarial networks with contrastive regularizers. CoRR, abs/1906.06034, 2019. URL http: //arxiv.org/abs/1906.06034
1906 arXiv
-
[48]
Patel, and Anima Anandkumar
Weili Nie, Tero Karras, Animesh Garg, Shoubhik Debnath, Anjul Pat- ney, Ankit B. Patel, and Anima Anandkumar. Semi-supervised style- gan for disentanglement learning. CoRR, abs/2003.03461, 2020. URL https://arxiv.org/abs/2003.03461
2003 arXiv
-
[50]
Towards building a group-based unsupervised representation disentanglement framework
Tao Yang, Xuanchi Ren, Yuwang Wang, Wenjun Zeng, and Nanning Zheng. Towards building a group-based unsupervised representation disentanglement framework. In ICLR, 2022
2022
-
[51]
Recur- sive disentanglement network
Yixuan Chen, Yubin Shi, Dongsheng Li, Yujiang Wang, Mingzhi Dong, Yingying Zhao, Robert Dick, Qin Lv, Fan Yang, and Li Shang. Recur- sive disentanglement network. In ICLR, 2022
2022
-
[52]
Learning disentangled representation by exploiting pretrained generative models: A contrastive learning view
Xuanchi Ren, Tao Yang, Yuwang Wang, and Wenjun Zeng. Learning disentangled representation by exploiting pretrained generative models: A contrastive learning view. In ICLR, 2022
2022
-
[53]
Multiplicative binding, representation operators, and anal- ogy
Ross Gayler. Multiplicative binding, representation operators, and anal- ogy. In Advances in Analogy Research: Integration of Theory and Data from the Cognitive, Computational, and Neural Sciences, pages 1–4, 01 1998
1998
-
[54]
Holographic reduced representations: Convolution alge- bra for compositional distributed representations
Tony Plate. Holographic reduced representations: Convolution alge- bra for compositional distributed representations. In Proceedings of the 12th International Joint Conference on Artificial Intelligence - Vol- ume 1, IJCAI’91, page 30–35, San Francisco, CA, USA, 1991. Morgan K...
1991
-
[55]
T. A. Plate. Holographic Reduced Representations: Distributed Rep- resentation for Cognitive Structures. Stanford: Center for the Study of Language and Information (CSLI), USA, 2003
2003
-
[56]
How to build a brain: A neural architecture for bio- logical cognition
Chris Eliasmith. How to build a brain: A neural architecture for bio- logical cognition. Oxford University Press, 2013
2013
-
[57]
A neural representation of continuous space using fractional binding
Brent Komer, Terrence C Stewart, Aaron R V oelker, and Chris Elia- smith. A neural representation of continuous space using fractional binding. In 41st annual meeting of the cognitive science society . QC: Cognitive Science Society, 2019
2019
-
[58]
Gayler and Simon D
Ross W. Gayler and Simon D. Levy. A distributed basis for analog- ical mapping, 2009. URL https://api.semanticscholar.org/CorpusID: 18842042
2009
-
[59]
Rachkovskij, Evgeny Osipov, and Jan M
Denis Kleyko, Abbas Rahimi, Dmitri A. Rachkovskij, Evgeny Osipov, and Jan M. Rabaey. Classification and recall with binary hyperdimen- sional computing: Tradeoffs in choice of density and mapping charac- teristics. IEEE Transactions on Neural Networks and Learning Systems, 29(...
2018
-
[60]
A comparison of vec- tor symbolic architectures
Kenny Schlegel, Peer Neubert, and Peter Protzel. A comparison of vec- tor symbolic architectures. Artificial Intelligence Review, 55(6):4523– 4555, 2022
2022
-
[61]
Multi- level variational autoencoder: Learning disentangled representations from grouped observations
Diane Bouchacourt, Ryota Tomioka, and Sebastian Nowozin. Multi- level variational autoencoder: Learning disentangled representations from grouped observations. In Proceedings of the AAAI Conference on Artificial Intelligence, 2018
2018
-
[62]
Disentangling factors of variation with cycle-consistent vari- ational auto-encoders
Ananya Harsh Jha, Saket Anand, Maneesh Singh, and VS Rao Veer- avasarapu. Disentangling factors of variation with cycle-consistent vari- ational auto-encoders. In Proceedings of the European Conference on Computer Vision (ECCV), pages 805–820, 2018
2018
-
[63]
Unsupervised robust disentangling of latent characteristics for image synthesis
Patrick Esser, Johannes Haux, and Bjorn Ommer. Unsupervised robust disentangling of latent characteristics for image synthesis. In Proceed- ings of the IEEE/CVF International Conference on Computer Vision , pages 2699–2709, 2019
2019
-
[64]
Learn- ing disentangled representations via mutual information estimation
Eduardo Hugo Sanchez, Mathieu Serrurier, and Mathias Ortner. Learn- ing disentangled representations via mutual information estimation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXII 16 , pages 205–221. Springer, 2020
2020
-
[65]
Nest- edvae: Isolating common factors via weak supervision
Matthew J V owels, Necati Cihan Camgoz, and Richard Bowden. Nest- edvae: Isolating common factors via weak supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, pages 9202–9212, 2020
2020
-
[66]
Towards a defi- nition of disentangled representations, 2018
Irina Higgins, David Amos, David Pfau, Sebastien Racaniere, Loic Matthey, Danilo Rezende, and Alexander Lerchner. Towards a defi- nition of disentangled representations, 2018
2018
-
[67]
Montero, Jeffrey S
Milton L. Montero, Jeffrey S. Bowers, Rui Ponte Costa, Casimir J. H. Ludwig, and Gaurav Malhotra. Lost in latent space: Disentangled models and the challenge of combinatorial generalisation, 2022. URL https://arxiv.org/abs/2204.02283
2022 arXiv
-
[68]
Decoupled weight decay regulariza- tion, 2017
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regulariza- tion, 2017. URL https://arxiv.org/abs/1711.05101
2017 arXiv
-
[69]
Smith and Nicholay Topin
Leslie N. Smith and Nicholay Topin. Super-convergence: Very fast training of neural networks using large learning rates, 2017. URL https://arxiv.org/abs/1708.07120
2017 arXiv
-
[70]
Deep Resid- ual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Resid- ual Learning for Image Recognition. In Proceedings of 2016 IEEE Conference on Computer Vision and Pattern Recognition , CVPR ’16, pages 770–778. IEEE, June 2016. doi: 10.1109/CVPR.2016.90. URL http://ieeexplore...
2016
-
[71]
Gans trained by a two time-scale up- date rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale up- date rule converge to a local nash equilibrium. In Proceedings of the 31st International Conference on Neural Information Processing Sys- tems, NIPS’...
2017
-
[72]
Bridging the gap to real-world object-centric learning
Maximilian Seitzer, Max Horn, Andrii Zadaianchuk, Dominik Ziet- low, Tianjun Xiao, Carl-Johann Simon-Gabriel, Tong He, Zheng Zhang, Bernhard Schölkopf, Thomas Brox, et al. Bridging the gap to real-world object-centric learning. In The Eleventh International Conference on Learn...
2022
-
[73]
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021
2021
-
[74]
Kent, Bruno A
Denis Kleyko, Mike Davies, Edward Paxon Frady, Pentti Kanerva, Spencer J. Kent, Bruno A. Olshausen, Evgeny Osipov, Jan M. Rabaey, Dmitri A. Rachkovskij, Abbas Rahimi, and Friedrich T. Sommer. Vec- tor symbolic architectures as a computing framework for emerging hardware. Proce...
2022
-
[2017]
URL https://openreview.net/forum?id=Sy2fzU9gl
-
[2018]
URL http://arxiv.org/abs/1812.04948
-
[9964]
URL https://doi.org/10.1007/ s12559-009-9009-8
doi: 10.1007/s12559-009-9009-8. URL https://doi.org/10.1007/ s12559-009-9009-8
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.