Pith. sign in

REVIEW 1 cited by

PixelLink: Detecting Scene Text via Instance Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1801.01315 v1 pith:XCMPNDDG submitted 2018-01-04 cs.CV

classification cs.CV
keywords textsegmentationinstanceregressionsceneboundinglocationmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Most state-of-the-art scene text detection algorithms are deep learning based methods that depend on bounding box regression and perform at least two kinds of predictions: text/non-text classification and location regression. Regression plays a key role in the acquisition of bounding boxes in these methods, but it is not indispensable because text/non-text prediction can also be considered as a kind of semantic segmentation that contains full location information in itself. However, text instances in scene images often lie very close to each other, making them very difficult to separate via semantic segmentation. Therefore, instance segmentation is needed to address this problem. In this paper, PixelLink, a novel scene text detection algorithm based on instance segmentation, is proposed. Text instances are first segmented out by linking pixels within the same instance together. Text bounding boxes are then extracted directly from the segmentation result without location regression. Experiments show that, compared with regression-based methods, PixelLink can achieve better or comparable performance on several benchmarks, while requiring many fewer training iterations and less training data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SAViL-Det: Semantic-Aware Vision-Language Model for Multi-Script Text Detection

    cs.CV 2025-07 reject novelty 4.0 of 10

    SAViL-Det combines CLIP, an asymptotic feature pyramid, and cross-modal attention to report F-scores of 84.8 on MLT-2019 and 90.2 on CTW1500, claiming state-of-the-art multi-script and curved text detection.

Pith tools