REVIEW 2 cited by
Transferring Rich Feature Hierarchies for Robust Visual Tracking
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Convolutional neural network (CNN) models have demonstrated great success in various computer vision tasks including image classification and object detection. However, some equally important tasks such as visual tracking remain relatively unexplored. We believe that a major hurdle that hinders the application of CNN to visual tracking is the lack of properly labeled training data. While existing applications that liberate the power of CNN often need an enormous amount of training data in the order of millions, visual tracking applications typically have only one labeled example in the first frame of each video. We address this research issue here by pre-training a CNN offline and then transferring the rich feature hierarchies learned to online tracking. The CNN is also fine-tuned during online tracking to adapt to the appearance of the tracked target specified in the first video frame. To fit the characteristics of object tracking, we first pre-train the CNN to recognize what is an object, and then propose to generate a probability map instead of producing a simple class label. Using two challenging open benchmarks for performance evaluation, our proposed tracker has demonstrated substantial improvement over other state-of-the-art trackers.
Forward citations
Cited by 2 Pith papers
-
Visual Object Tracking across Diverse Data Modalities: A Review
A survey that taxonomizes deep-learning visual trackers across RGB, thermal, LiDAR, and four multi-modal combinations, with benchmark tables and future directions.
-
Object Tracking in a $360^o$ View: A Novel Perspective on Bridging the Gap to Biomedical Advancements
A review of object tracking algorithms for biomedical video concludes that deep learning is the most capable family, but it includes a placeholder citation for a model described as real.
Discussion (0). Continue with ORCID to comment.