Pith. sign in

Self-taught Object Localization with Deep Networks

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

This paper introduces self-taught object localization, a novel approach that leverages deep convolutional networks trained for whole-image recognition to localize objects in images without additional human supervision, i.e., without using any ground-truth bounding boxes for training. The key idea is to analyze the change in the recognition scores when artificially masking out different regions of the image. The masking out of a region that includes the object typically causes a significant drop in recognition score. This idea is embedded into an agglomerative clustering technique that generates self-taught localization hypotheses. Our object localization scheme outperforms existing proposal methods in both precision and recall for small number of subwindow proposals (e.g., on ILSVRC-2012 it produces a relative gain of 23.4% over the state-of-the-art for top-1 hypothesis). Furthermore, our experiments show that the annotations automatically-generated by our method can be used to train object detectors yielding recognition results remarkably close to those obtained by training on manually-annotated bounding boxes.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2019 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

Learning Rich Representations For Structured Visual Prediction Tasks

cs.CV · 2019-08-30 · conditional · novelty 4.0

Nested multi-scale 'zoom-out' features around each pixel or superpixel give competitive semantic segmentation with a simple feedforward classifier, and segmentation maps act as surprisingly strong inputs for depth prediction.

citing papers explorer

Showing 1 of 1 citing paper.

  • Learning Rich Representations For Structured Visual Prediction Tasks cs.CV · 2019-08-30 · conditional · none · ref 96 · internal anchor

    Nested multi-scale 'zoom-out' features around each pixel or superpixel give competitive semantic segmentation with a simple feedforward classifier, and segmentation maps act as surprisingly strong inputs for depth prediction.