REVIEW 22 cited by
On the Binding Problem in Artificial Neural Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
On the Binding Problem in Artificial Neural Networks
read the original abstract
Contemporary neural networks still fall short of human-level generalization, which extends far beyond our direct experiences. In this paper, we argue that the underlying cause for this shortcoming is their inability to dynamically and flexibly bind information that is distributed throughout the network. This binding problem affects their capacity to acquire a compositional understanding of the world in terms of symbol-like entities (like objects), which is crucial for generalizing in predictable and systematic ways. To address this issue, we propose a unifying framework that revolves around forming meaningful entities from unstructured sensory inputs (segregation), maintaining this separation of information at a representational level (representation), and using these entities to construct new inferences, predictions, and behaviors (composition). Our analysis draws inspiration from a wealth of research in neuroscience and cognitive psychology, and surveys relevant mechanisms from the machine learning literature, to help identify a combination of inductive biases that allow symbolic information processing to emerge naturally in neural networks. We believe that a compositional approach to AI, in terms of grounded symbol-like representations, is of fundamental importance for realizing human-level generalization, and we hope that this paper may contribute towards that goal as a reference and inspiration.
Forward citations
Cited by 22 Pith papers
-
How can embedding models bind concepts?
CLIP relies on high-complexity additive binding that prevents generalization to unseen concept combinations, whereas transformers trained from scratch develop low-complexity multiplicative binding functions that enabl...
-
SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding
SYNCR benchmark shows leading MLLMs reach only 52.5% average accuracy on cross-video reasoning tasks against an 89.5% human baseline, with major weaknesses in physical and spatial reasoning.
-
Learning to Theorize the World from Observation
NEO is a probabilistic neural model that induces compositional programs as a learned Language of Thought from non-textual observations and executes them via a shared transition model to enable explanation-driven gener...
-
ActionParty: Multi-Subject Action Binding in Generative Video Games
ActionParty binds discrete actions to individual subjects in a single generated video by jointly modeling subject state tokens and video latents, controlling up to seven players across 46 Melting Pot games.
-
Mechanistic Independence: A Principle for Identifiable Disentangled Representations
Mechanistic independence criteria yield identifiability of latent subspaces under nonlinear mixing by focusing on action-based independence rather than latent distributions, with a hierarchy and graph-theoretic view o...
-
Human-like Object Grouping in Self-supervised Vision Transformers
DINO self-supervised transformers best predict human same/different object RTs; object-centric patch affinity and Gram-matrix distillation explain and transfer the alignment.
-
EditCLEVR: A Paired-Scene Intervention Benchmark for Compositional Faithfulness of Object-Centric Representations
EditCLEVR uses before/after scene pairs with one known object edit to show that object-centric models fail to decode exactly the intended single-object attribute change under a CoGenT color-shape shift, even with grou...
-
HSA: Hierarchical Slot Attention for Multi-granularity Scene-Decomposition
One hierarchical slot-attention model with 10% labels jointly yields holistic, semantic, and panoptic scene decompositions that outperform three separate flat baselines by large ARI margins.
-
Input Pathways Shape Few-Shot, Not Zero-Shot, Binding in Tiny Transformers: A Fully-Enumerable Study
In information-matched tiny transformers, zero-shot compositional binding fails for every route, while few-shot efficiency is governed by input-pathway sharing and code readability.
-
Dual-State Slot Attention: Decoupling Appearance and Identity for Video Object-Centric Learning
DSSA decouples per-frame appearance from temporal identity in slot attention mechanisms to reduce slot swapping and improve temporal consistency in video object segmentation.
-
No Epoch Like the Present: Robust Climate Emulation Requires Out-of-Distribution Generalisation
ML climate emulators degrade under seasonal distribution shifts that proxy long-term climate change, but physically motivated compositional decompositions improve out-of-distribution performance with modest in-distrib...
-
Slot-MPC: Goal-Conditioned Model Predictive Control with Object-Centric Representations
Slot-MPC learns slot representations to build a differentiable object-centric dynamics model that supports efficient gradient-based MPC for robotic manipulation in novel situations.
-
Learning to Theorize the World from Observation
NEO induces compositional latent programs as world theories from observations and executes them to enable explanation-driven generalization.
-
Deep Sprite-based Image Models: An Analysis
A deep sprite-based image decomposition method matches SOTA unsupervised class-aware segmentation on CLEVR, scales linearly with objects, explicitly identifies categories, and fully models images interpretably.
-
I Walk the Line: Examining the Role of Gestalt Continuity in Object Binding for Vision Transformers
Pretrained vision transformers use specific attention heads sensitive to Gestalt continuity for object binding, shown via probes on synthetic datasets and ablation experiments.
-
Kuramoto Oscillatory Phase Encoding: Neuro-inspired Synchronization for Improved Learning Efficiency
Kuramoto oscillatory phase encoding (KoPE) adds evolving phase states to Vision Transformers and claims better training, parameter, and data efficiency via synchronization-enhanced structure learning.
-
ORGAN: Object-Centric Representation Learning using Cycle Consistent Generative Adversarial Networks
A cycle-consistent GAN that translates between images and object lists matches state-of-the-art detection on synthetic scenes and detects low-contrast cells where slot-attention models fail.
-
From Cortical Synchronous Rhythm to Brain Inspired Learning Mechanism: An Oscillatory Spiking Neural Network with Time-Delayed Coordination
S2-Net is an oscillatory spiking neural network that uses time-delayed synchronization for bottom-up and top-down coordination to enable efficient, brain-inspired information processing across tasks like decoding and ...
-
Deep Sprite-based Image Models: An Analysis
A deep sprite-based image model, guided by a design analysis, matches SOTA unsupervised class-aware segmentation on CLEVR, scales linearly with objects, and yields fully interpretable decompositions.
-
Kuramoto Oscillatory Phase Encoding: Neuro-inspired Synchronization for Improved Learning Efficiency
KoPE adds Kuramoto-based oscillatory phase states and synchronization to Vision Transformers, improving training, parameter, and data efficiency on structured vision tasks.
-
Dynamics Reveals Structure: Challenging the Linear Propagation Assumption
Under the linear propagation assumption, first-order feature geometry cannot support both negation and composition: the only feature map satisfying both is the zero map.
-
Structural Prognostic Event Modeling for Multimodal Cancer Survival Analysis
SlotSPE is a slot-attention framework that decomposes multimodal cancer data into structural prognostic event slots to improve survival prediction and interpretability.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.