Pith. sign in

REVIEW 1 cited by

Self-supervised Visual Reinforcement Learning with Object-centric Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.14381 v1 pith:DCBCXX3W submitted 2020-11-29 cs.LG

Self-supervised Visual Reinforcement Learning with Object-centric Representations

classification cs.LG
keywords skillsagentautonomouscompositionalrepresentationsdiscoverobject-centricscene
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Autonomous agents need large repertoires of skills to act reasonably on new tasks that they have not seen before. However, acquiring these skills using only a stream of high-dimensional, unstructured, and unlabeled observations is a tricky challenge for any autonomous agent. Previous methods have used variational autoencoders to encode a scene into a low-dimensional vector that can be used as a goal for an agent to discover new skills. Nevertheless, in compositional/multi-object environments it is difficult to disentangle all the factors of variation into such a fixed-length representation of the whole scene. We propose to use object-centric representations as a modular and structured observation space, which is learned with a compositional generative world model. We show that the structure in the representations in combination with goal-conditioned attention policies helps the autonomous agent to discover and learn useful skills. These skills can be further combined to address compositional tasks like the manipulation of several different objects.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ES-Merging: Biological MLLM Merging via Embedding Space Signals

    cs.LG 2026-03 unverdicted novelty 5.0

    ES-Merging estimates layer-wise and element-wise merge coefficients from coarse- and fine-grained embedding signals and claims better cross-modal reasoning and single-modal knowledge preservation than parameter-space merging.