REVIEW 8 cited by
Object-Centric Learning with Slot Attention
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Learning object-centric representations of complex scenes is a promising step towards enabling efficient abstract reasoning from low-level perceptual features. Yet, most deep learning approaches learn distributed representations that do not capture the compositional properties of natural scenes. In this paper, we present the Slot Attention module, an architectural component that interfaces with perceptual representations such as the output of a convolutional neural network and produces a set of task-dependent abstract representations which we call slots. These slots are exchangeable and can bind to any object in the input by specializing through a competitive procedure over multiple rounds of attention. We empirically demonstrate that Slot Attention can extract object-centric representations that enable generalization to unseen compositions when trained on unsupervised object discovery and supervised property prediction tasks.
Forward citations
Cited by 8 Pith papers
-
Grounding Spatial Relations in a Compact World Model: Instruction Leakage and a Goal-Free Dynamics Fix
Goal-conditioned world models transcribe instructions instead of perceiving spatial relations when the instruction names the scored quantity, and removing the goal from the dynamics fixes it.
-
Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties
ViWi uses material slots and a simulated RF descriptor to predict voxel-level Young's modulus, Poisson's ratio, and density, reporting gains over prior work on a synthetic benchmark.
-
Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models
ORCA splits Q-Former queries into orthogonally constrained groups, reversing directional collapse and speaker-indistinguishability in audio-LLM connectors and gaining 26.4 points on SAKURA multi-hop reasoning.
-
Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation
Slot-based object-centric visual representations, especially with robot-video pretraining, improve out-of-distribution generalization of robotic manipulation policies compared to global and dense pre-trained features.
-
Successes and Limitations of Object-centric Models at Compositional Generalisation
Object-centric models handle novel combinations of object properties when all local features are present, and can even extrapolate to unseen shapes on the Pentomino dataset.
-
Mixing Configurations for Downstream Prediction
Mixing multiple resolution clusterings of an embedding, aligned between train and test, and fused by attention, improves downstream regression and classification over single-resolution baselines.
-
Agent Identity Evals: Measuring Agentic Identity
Introduces Agent Identity Evals (AIE), five similarity-based metrics for LMA identity stability, with pilot experiments showing identifiability always at zero and no statistical support.
-
Is an object-centric representation beneficial for robotic manipulation ?
Evaluating the object-centric SAVi encoder against the global DINO and R3M representations on three simulated manipulation tasks, the authors find SAVi is the only model to solve the pick task and is more robust to un...
Discussion (0). Continue with ORCID to comment.