Pith. sign in

REVIEW 5 cited by

R2CNN: Rotational Region CNN for Orientation Robust Scene Text Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1706.09579 v2 pith:3M6HBF26 submitted 2017-06-29 cs.CV

classification cs.CV
keywords textaxis-aligneddetectionregiondifferentfeaturesicdarinclined
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we propose a novel method called Rotational Region CNN (R2CNN) for detecting arbitrary-oriented texts in natural scene images. The framework is based on Faster R-CNN [1] architecture. First, we use the Region Proposal Network (RPN) to generate axis-aligned bounding boxes that enclose the texts with different orientations. Second, for each axis-aligned text box proposed by RPN, we extract its pooled features with different pooled sizes and the concatenated features are used to simultaneously predict the text/non-text score, axis-aligned box and inclined minimum area box. At last, we use an inclined non-maximum suppression to get the detection results. Our approach achieves competitive results on text detection benchmarks: ICDAR 2015 and ICDAR 2013.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Unconstrained End-to-End Text Spotting

    cs.CV 2019-08 conditional novelty 7.0 of 10

    A Mask R-CNN and attention-based text spotter handles curved text by masking RoI features instead of rectifying them, and uses OCR-engine labels as extra training data to set state-of-the-art results on ICDAR15 and To...

  2. Bridging Severe Cross-Modal Misalignment: End-to-End Visible-Infrared Object Detection via Explicit Feature-Domain Affine Registration

    cs.CV 2026-08 conditional novelty 6.0 of 10

    JFRDet aligns infrared and visible features with an explicit affine transformation and illumination-guided fusion, reaching 69.7% mAP50 on the new DVMA benchmark.

  3. FIction: 4D Future Interaction Prediction from Video

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FICTION predicts future 3D interaction locations and body poses up to three minutes ahead from egocentric video and a 3D scene map, and claims substantial gains over prior methods on a new Ego-Exo4D benchmark.

  4. Geometry Normalization Networks for Accurate Scene Text Detection

    cs.CV 2019-09 conditional novelty 6.0 of 10

    Scene text detectors improve by routing feature maps through parallel branches that normalize scale and orientation before a shared detection head.

  5. R3Det: Refined Single-Stage Detector with Feature Refinement for Rotating Object

    cs.CV 2019-08 conditional novelty 6.0 of 10

    R3Det improves single-stage rotated-object detection through progressive horizontal-to-rotated refinement, a feature refinement module that realigns features by interpolation, and a SkewIoU-weighted regression loss.

Pith tools