REVIEW 9 cited by
CenterNet: Keypoint Triplets for Object Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In object detection, keypoint-based approaches often suffer a large number of incorrect object bounding boxes, arguably due to the lack of an additional look into the cropped regions. This paper presents an efficient solution which explores the visual patterns within each cropped region with minimal costs. We build our framework upon a representative one-stage keypoint-based detector named CornerNet. Our approach, named CenterNet, detects each object as a triplet, rather than a pair, of keypoints, which improves both precision and recall. Accordingly, we design two customized modules named cascade corner pooling and center pooling, which play the roles of enriching information collected by both top-left and bottom-right corners and providing more recognizable information at the central regions, respectively. On the MS-COCO dataset, CenterNet achieves an AP of 47.0%, which outperforms all existing one-stage detectors by at least 4.9%. Meanwhile, with a faster inference speed, CenterNet demonstrates quite comparable performance to the top-ranked two-stage detectors. Code is available at https://github.com/Duankaiwen/CenterNet.
Forward citations
Cited by 9 Pith papers
-
Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation
Screen2AX generates hierarchical macOS accessibility metadata from a screenshot and reports improved GPT-4 UI task success compared with native accessibility and OmniParser V2.
-
Training-Time-Friendly Network for Real-Time Object Detection
TTFNet is an object detector that trains much faster than prior real-time detectors while keeping similar accuracy, by using Gaussian-weighted samples around each object center during training.
-
Deep High-Resolution Representation Learning for Visual Recognition
HRNet maintains high-resolution feature maps in parallel with low-resolution streams and repeatedly fuses them, improving accuracy on pose, segmentation, detection, and face alignment benchmarks.
-
Seq-SG2SL: Inferring Semantic Layout from Scene Graph Through Sequence to Sequence Learning
A Transformer sequence-to-sequence model translates scene graphs into semantic layouts via new token codes, and a new metric SLEU measures layout similarity.
-
Matrix Nets: A New Deep Architecture for Object Detection
The paper introduces Matrix Nets, a feature-pyramid-like architecture with separate layers for scale and aspect ratio, and shows a keypoint-based detector built on it reaches 47.8 mAP on MS COCO.
-
Revisiting Feature Alignment for One-stage Object Detection
A one-stage detector that aligns convolutional features to predicted anchor boxes via a RoIConv operator, achieving 44.1 mAP on COCO test-dev.
-
AFP-Net: Realtime Anchor-Free Polyp Detection in Colonoscopy
AFP-Net, an anchor-free polyp detector with a context enhancement module and cosine ground-truth projection, achieves 99.36% precision and 96.44% recall on CVC-Clinic, and 52.6 FPS.
-
SHeRL-FL: When Representation Learning Meets Split Learning in Hierarchical Federated Learning
The submitted body is an unrelated survey, not the SHeRL-FL method claimed in the metadata.
-
Recent Advances in Deep Learning for Object Detection
A structured survey of deep learning object detection covering two-stage and one-stage detectors, feature learning, training strategies, applications, and benchmarks up to 2019.
Discussion (0). Continue with ORCID to comment.