Pith. sign in

super hub

Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

35 Pith papers cite this work, alongside 18,238 external citations. Polarity classification is still indexing.

35 Pith papers citing it
18.2k external citations · Pith
abstract

State-of-the-art object detection networks depend on region proposal algorithms to hypothesize object locations. Advances like SPPnet and Fast R-CNN have reduced the running time of these detection networks, exposing region proposal computation as a bottleneck. In this work, we introduce a Region Proposal Network (RPN) that shares full-image convolutional features with the detection network, thus enabling nearly cost-free region proposals. An RPN is a fully convolutional network that simultaneously predicts object bounds and objectness scores at each position. The RPN is trained end-to-end to generate high-quality region proposals, which are used by Fast R-CNN for detection. We further merge RPN and Fast R-CNN into a single network by sharing their convolutional features---using the recently popular terminology of neural networks with 'attention' mechanisms, the RPN component tells the unified network where to look. For the very deep VGG-16 model, our detection system has a frame rate of 5fps (including all steps) on a GPU, while achieving state-of-the-art object detection accuracy on PASCAL VOC 2007, 2012, and MS COCO datasets with only 300 proposals per image. In ILSVRC and COCO 2015 competitions, Faster R-CNN and RPN are the foundations of the 1st-place winning entries in several tracks. Code has been made publicly available.

hub tools

citation-role summary

background 2 dataset 1 method 1

citation-polarity summary

representative citing papers

Tri-Modal Fusion Transformers for UAV-based Object Detection

cs.CV · 2026-04-17 · unverdicted · novelty 7.0

A dual-stream vision transformer with modality-aware gated exchange and bidirectional token exchange fuses RGB, thermal, and event data to improve UAV vehicle detection over dual-modal baselines on a new 10,489-frame dataset.

RefDiffNet: Learning to Expose Subtle PCB Defects Before Detection

cs.CV · 2026-05-30 · unverdicted · novelty 6.0

RefDiffNet is a lightweight input enhancement block that uses reference image comparison to expose PCB defects, delivering up to 18% relative mAP50:95 gains across YOLO, RT-DETR, and Faster R-CNN detectors with 0.004-0.005M extra parameters.

New VVC profiles targeting Feature Coding for Machines

cs.CV · 2025-12-09 · unverdicted · novelty 4.0

Three lightweight VVC profiles for feature coding achieve up to 2.96% BD-Rate gain and 95.6% encoding speedup while preserving downstream task accuracy under the MPEG-AI FCM framework.

citing papers explorer

Showing 35 of 35 citing papers.