Pith. sign in

REVIEW 1 cited by

Self-supervised visual learning from interactions with objects

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.06704 v2 pith:K7WFWCAU submitted 2024-07-09 cs.CV cs.LG

classification cs.CVcs.LG
keywords learningobjectsvisualactionsinteractionsobjectperformedaction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised learning (SSL) has revolutionized visual representation learning, but has not achieved the robustness of human vision. A reason for this could be that SSL does not leverage all the data available to humans during learning. When learning about an object, humans often purposefully turn or move around objects and research suggests that these interactions can substantially enhance their learning. Here we explore whether such object-related actions can boost SSL. For this, we extract the actions performed to change from one ego-centric view of an object to another in four video datasets. We then introduce a new loss function to learn visual and action embeddings by aligning the performed action with the representations of two images extracted from the same clip. This permits the performed actions to structure the latent visual representation. Our experiments show that our method consistently outperforms previous methods on downstream category recognition. In our analysis, we find that the observed improvement is associated with a better viewpoint-wise alignment of different objects from the same category. Overall, our work demonstrates that embodied interactions with objects can improve SSL of object categories.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Human Gaze Boosts Object-Centered Representation Learning

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Gaze-centered cropping of egocentric video improves self-supervised object representation learning, especially fine-grained and instance recognition.

Pith tools