Pith. sign in

REVIEW 1 cited by

Learning to Navigate Using Mid-Level Visual Priors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.11121 v1 pith:DQJJE5D3 submitted 2019-12-23 cs.CV cs.LGcs.NEcs.RO

classification cs.CVcs.LGcs.NEcs.RO
keywords learningvisualdownstreammid-levelpriorstasksvisionworld
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How much does having visual priors about the world (e.g. the fact that the world is 3D) assist in learning to perform downstream motor tasks (e.g. navigating a complex environment)? What are the consequences of not utilizing such visual priors in learning? We study these questions by integrating a generic perceptual skill set (a distance estimator, an edge detector, etc.) within a reinforcement learning framework (see Fig. 1). This skill set ("mid-level vision") provides the policy with a more processed state of the world compared to raw images. Our large-scale study demonstrates that using mid-level vision results in policies that learn faster, generalize better, and achieve higher final performance, when compared to learning from scratch and/or using state-of-the-art visual and non-visual representation learning methods. We show that conventional computer vision objectives are particularly effective in this regard and can be conveniently integrated into reinforcement learning frameworks. Finally, we found that no single visual representation was universally useful for all downstream tasks, hence we computationally derive a task-agnostic set of representations optimized to support arbitrary downstream tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Object-Centric Action-Enhanced Representations for Robot Visuo-Motor Policy Learning

    cs.RO 2025-05 conditional novelty 5.0 of 10

    An object-centric encoder fine-tuned on human action videos improves simulated robot pouring policies in behavioral cloning and offline reinforcement learning.

Pith tools