REVIEW 3 major objections 7 minor 57 references
Hyper-Local Deformable Transformers for Text Spotting on Historical Maps
T0 review · 3 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A text spotting model that anchors deformable attention to each predicted character center and boundary point outperforms state-of-the-art spotters on historical maps, with the largest gains on long and highly rotated labels.
desk verdict Solid applied paper with a real mechanism behind the gains; the iterative character-center training deserves a skeptical read but doesn't sink the central claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is hyper-local sampling inside a Deformable DETR-style decoder. For each content query representing a boundary point or a character, the reference point for deformable attention is the query's own predicted location rather than the text-instance center, and the network samples image features around that location with learned offsets. A character-center predictor computes character centers from character and boundary-point queries via cross-attention, and the predicted positions are also injected as hyper-local positional embeddings so each sub-component knows where it sits relative to the others. This gives every query an explicit local positional prior and lets intra- and inter-instance self-attention reason about the spatial arrangement of boundary points and characters.
What would settle it
Rerun the paper's without-hyper-local-character-sampling ablation—fixing all character reference points at the proposal center—on the larger historical map benchmark and verify that end-to-end F-score drops by the reported roughly 5 points and that the drop concentrates on text of length 7 or more and on rotations above 60 degrees. If fixed-center sampling with identical training data matches or beats hyper-local sampling, the mechanism is not what carries the improvement. A second check: jitter the supervised character centers by a few pixels during training and measure whether end-to-end performance degrades, which would confirm the model actually depends on precise reference points.
Extended reading notes
Core claim
PaLeTTe establishes that hyper-local reference points—predicted character centers for recognition and predicted boundary points for detection—carry the positional information that deformable attention needs, and that coarse instance-level reference points are the main bottleneck. Starting from box-proposal centers, the decoder predicts boundary points, then a character-center predictor reads character and boundary queries to estimate character centers; in subsequent layers the sampling and positional embeddings are re-anchored to these refined predictions. On two new hand-annotated historical map benchmarks, this yields end-to-end spotting F-score gains of roughly 3 to 6 points over the strongest baselines, with the largest improvements concentrated on text longer than ten characters and on text rotated into the 60–90 degree range. The paper further shows that character-center supervision can be bootstrapped from boundary-only annotations through iterative training, which is what makes the method applicable to real maps where character centers are not labeled.
Load-bearing premise
The load-bearing premise is that the character-center predictor can learn accurate character centers from boundary points and character queries, and that any predicted center landing inside the ground-truth boundary is considered correct during iterative training; if those centers are systematically biased, hyper-local sampling anchors to the wrong image locations and the reported gains should shrink or disappear.
Editorial extensions
If this is right
- Long and highly rotated text, the cases where fixed-center sampling is most misaligned, show the largest gains: on text of length 10 or more the end-to-end recall rises by about 8 points on one benchmark, and for rotations of 60–90 degrees it rises by about 4.8 points on the larger benchmark.
- Adding SynthMap+ to pretraining improves PaLeTTe's end-to-end F-score by about 24.8 points on the larger benchmark and also improves the TESTR baseline, indicating the synthetic data is a reusable training resource rather than a model-specific fix.
- Ablating hyper-local sampling for boundary points costs about 2.9 points of detection F1, and ablating it for characters costs about 5.1 points of end-to-end performance, evidence that re-anchoring to predicted sub-component locations is the load-bearing design choice.
- The iterative training procedure lets human annotations without character-center labels be folded into training, which is what allows the model to be finetuned on real maps at scale.
- Deployed on more than 60,000 maps, the model generated over 100 million text labels used for full-text map search, an existence proof that the approach scales beyond benchmarks.
Reading between the lines
- If hyper-local re-anchoring is the cause of the gains, a testable extension is to apply the same decoder modification to scene-text spotters; a similar improvement there would show the mechanism is about sub-component alignment, not about the map domain.
- The SynthMap+ recipe—separating geometric text placement from background style and sourcing backgrounds from real scans—could be pointed at other document types, such as manuscripts or engineering drawings, to generate synthetic training data with the same ease.
- The failure cases the paper itself reports (very large characters, overlapping text and line features) suggest the method inherits a sensitivity to reference-point accuracy: when the character-center predictor is confidently wrong, hyper-local sampling could lock onto the wrong features, so an uncertainty estimate on predicted centers might extend the approach to its own hard cases.
- A further consequence left implicit is that the character-center predictor gives the model a free alignment signal between detection and recognition; one could exploit that alignment as a confidence measure for spotting, scoring a detection as more reliable when boundary points and character centers agree.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes PaLeTTe, an end-to-end text spotter for historical maps built on Deformable DETR. The core idea is hyper-local sampling: predicted boundary-point locations and character centers are used as reference points for deformable attention, replacing the coarse instance-center reference used in prior DETR-based spotters. A character-center predictor infers character centers from boundary points and character content queries, and an iterative training procedure adds predicted centers as pseudo-labels when they fall inside the ground-truth text boundary. The paper also introduces SynthMap+, a synthetic map-image generator, and a new annotated benchmark, Rumsey-309. Experiments on Rumsey-309 and Grinnell-UMass-31 show that PaLeTTe outperforms ABCNet-v2, SWINTS, TESTR, and DeepSolo, with ablations indicating that hyper-local sampling and hyper-local positional embeddings contribute to the gains.
Significance. If the empirical results hold, this is a practically valuable contribution: the deployment on 60,000 David Rumsey maps with over 100 million extracted labels demonstrates real-world impact. The release of code, synthetic data, and a new benchmark is a concrete strength, as is the standard training and evaluation protocol. The central mechanistic claim—that sampling features around predicted sub-component locations improves spotting on long and rotated map text—is well motivated and supported by the pretrained (non-finetuned) results in Table 1, which show the largest gains and do not rely on the iterative pseudo-label procedure. The main uncertainty is the quality and possible circularity of the pseudo-labeled character centers used in finetuning, which deserves a more careful analysis than the paper currently provides.
major comments (3)
- [§2.5] The criterion for accepting a predicted character center as 'correct' is that it falls inside the ground-truth text boundary. For long, curved, or highly rotated text, the polygon encloses a large area, so a center can be inside the boundary yet far from the true character position. Because these accepted centers are then used as training labels in subsequent iterations, the model can reinforce its own biased predictions rather than learning from independent supervision. The pretrained results in Table 1 do not use this procedure and are therefore unaffected, but the finetuned results and the claim that iterative training alleviates the need for character-center annotation depend on it. Please provide a quantitative evaluation of pseudo-label quality (e.g., distance of predicted centers to manually annotated centers on a held-out subset of Total-Text or the evaluation maps) and an ablation comparing finetuning with and without iterative training on both benchmarks, reporting E2E and detection F1. This would help determine whether the finetuned gains are robust or partly an artifact of self-training.
- [§4.4.5] The ablation results in Table 5 are not reported consistently with the text. The text states that removing hyper-local sampling gives a 2.9% reduction in detection F1 (first row) and a 5.1% reduction in recognition performance (second row), but Table 5 shows PaLeTTe-wo-HLD at 82.3 vs 84.7 detection F1 (a 2.4-point difference) and PaLeTTe-wo-HLR at 66.4 vs 69.8 E2E-None (a 3.4-point difference). Please correct the numbers and clarify whether the reductions are percentage points or relative percentages, and whether they refer to E2E-None or a separate recognition metric. In addition, please state explicitly whether the ablations are on the pretrained or finetuned models, since the paper does not specify this for Table 5.
- [§4.2.2] The newly introduced Rumsey-309 benchmark is central to the evaluation, but no annotation-quality statistics are reported. Please provide details on the annotation protocol: the number of annotators, the annotation tool, inter-annotator agreement on a subset, and any quality-control steps. It would also be useful to clarify whether any model selection or hyperparameter tuning was performed on the evaluation sets, since the authors are also the creators of the benchmark and the developers of the deployment system.
minor comments (7)
- [Figure 1 caption] The word 'Figrue' is misspelled as 'Figrue' in the caption.
- [Abstract and Introduction] The name 'PaLeTTe' is sometimes run together with adjacent words, e.g., 'PaLeTTewithSynthMap+' in the abstract and 'PaLeTTeprogressively' in Section 2. Please add spaces after 'PaLeTTe'.
- [§6.1] There is a missing space in 'Weinman et al. [50]propose' in the first paragraph.
- [§7] The benchmark name is written as 'Rumsey309' in the limitations section, while the rest of the paper uses 'Rumsey-309'. Please make the naming consistent.
- [Table 3] The orientation intervals are inconsistent: the Grinnell columns use [30, 60) and [60, 90], while the Rumsey columns use (30, 60] and (60, 90]. Please use consistent interval notation.
- [§2.2] The reference 'In Figure 2 1' should be 'In Figure 2(a)' or similar for clarity; the current notation appears to be a rendering artifact.
- [§2.3] In the character-center predictor equations, the dot notation in 'wq·(...)' is not standard for matrix-vector products; please use explicit linear projection notation for readability.
Circularity Check
No significant circularity: the paper's central claims are empirical and rest on external benchmarks and ablations; the iterative self-training loop and a contextual self-citation are not load-bearing.
full rationale
The central claim is empirical rather than derivational: PaLeTTe's hyper-local sampling and SynthMap+ are validated by ablations (Table 5) and by comparisons with ABCNet-v2, SWINTS, TESTR, and DeepSolo on Grinnell-UMass-31 and Rumsey-309 (Tables 1-4). The hyper-local reference points are predicted boundary points and character centers, but the reported metric is end-to-end spotting against independent ground-truth boundaries and transcriptions, so the mechanism is not defined in terms of the outcome. Section 2.5 iterative training is a self-training heuristic, not a circular derivation: predicted centers are added only when all centers lie inside the GT boundary, and the paper explicitly acknowledges that imperfect center localization limits training (Section 4.4.1, Section 7); the final evaluation does not reward center accuracy. The only notable self-citation, mapKurator [21] in Section 5, is deployment context and is not load-bearing for the SOTA claim. No equation reduces a prediction to a fitted parameter, and no uniqueness theorem or ansatz is imported from the authors' prior work. Possible concerns about test-set-based finetuning-dataset selection or benchmark construction are correctness risks, not circularity.
Assumptions & free parameters
free parameters (7)
- lambda_cls
- lambda_coord
- lambda_center
- lambda_ct
- lambda_char
- num_boundary_points =
16
- max_text_length =
25
assumptions (5)
- domain assumption Character centers are strongly related to character content and text instance boundary, allowing prediction from boundary points and character queries.
- domain assumption Text instances in historical maps can be adequately modeled by boundary point polygons and character centers.
- domain assumption Fine-tuning on Total-Text transfers to historical map text.
- domain assumption SynthMap+ backgrounds and text placement simulate the distribution of the evaluation datasets.
- standard math The Hungarian algorithm with the proposed cost produces a valid matching for training.
Cite this review
Pith. "Pith review of Hyper-Local Deformable Transformers for Text Spotting on Historical Maps." pith.science (2026). https://pith.science/paper/PC5DCJ5W
@misc{pith2026250615010,
author = {Pith},
title = {Pith review of: Hyper-Local Deformable Transformers for Text Spotting on Historical Maps},
year = {2026},
howpublished = {\url{https://pith.science/paper/PC5DCJ5W}},
note = {Machine review of arXiv:2506.15010}
}
read the original abstract
Text on historical maps contains valuable information providing georeferenced historical, political, and cultural contexts. However, text extraction from historical maps is challenging due to the lack of (1) effective methods and (2) training data. Previous approaches use ad-hoc steps tailored to only specific map styles. Recent machine learning-based text spotters (e.g., for scene images) have the potential to solve these challenges because of their flexibility in supporting various types of text instances. However, these methods remain challenges in extracting precise image features for predicting every sub-component (boundary points and characters) in a text instance. This is critical because map text can be lengthy and highly rotated with complex backgrounds, posing difficulties in detecting relevant image features from a rough text region. This paper proposes PALETTE, an end-to-end text spotter for scanned historical maps of a wide variety. PALETTE introduces a novel hyper-local sampling module to explicitly learn localized image features around the target boundary points and characters of a text instance for detection and recognition. PALETTE also enables hyper-local positional embeddings to learn spatial interactions between boundary points and characters within and across text instances. In addition, this paper presents a novel approach to automatically generate synthetic map images, SynthMap+, for training text spotters for historical maps. The experiment shows that PALETTE with SynthMap+ outperforms SOTA text spotters on two new benchmark datasets of historical maps, particularly for long and angled text. We have deployed PALETTE with SynthMap+ to process over 60,000 maps in the David Rumsey Historical Map collection and generated over 100 million text labels to support map searching. The project is released at https://github.com/kartta-foundation/mapkurator-palette-doc.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Samantha T Arundel, Trenton P Morgan, and Phillip T Thiem. 2022. Deep Learning Detection and Recognition of Spot Elevations on Historical Topographic Maps.Frontiers in Environmental Science(2022), 117
work page 2022
-
[2]
2017.A data set of annotated historical maps
Larry Boateng Asante, David Cambronero Sanchez, Ravi Chande, Dylan Gumm, and Jerod Weinman. 2017.A data set of annotated historical maps. Technical Report. Technical report, Grinnell College, Grinnell, Iowa
work page 2017
-
[3]
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexan- der Kirillov, and Sergey Zagoruyko. 2020. End-to-end object detection with transformers. InEuropean conference on computer vision. Springer, 213–229
2020
-
[4]
Yao-Yi Chiang, Muhao Chen, Weiwei Duan, Jina Kim, Craig A Knoblock, Stefan Leyk, Zekun Li, Yijun Lin, Min Namgung, Basel Shbita, et al. 2023. GeoAI for the Digitization of Historical Maps. InHandbook of Geospatial Artificial Intelligence. CRC Press, 217–247
work page 2023
-
[5]
2020.Using historical maps in scientific studies: Applications, challenges, and best practices
Yao-Yi Chiang, Weiwei Duan, Stefan Leyk, Johannes H Uhl, and Craig A Knoblock. 2020.Using historical maps in scientific studies: Applications, challenges, and best practices. Springer
work page 2020
-
[6]
Yao-Yi Chiang and Craig A. Knoblock. 2011. Recognition of Multi-oriented, Multi-sized, and Curved Text. In2011 International Conference on Document Analysis and Recognition. 1399–1403. https://doi.org/10.1109/ICDAR.2011.281 Yijun Lin & Yao-Yi Chiang
-
[7]
Yao-Yi Chiang, Stefan Leyk, and Craig A Knoblock. 2014. A survey of digital map processing techniques.ACM Computing Surveys (CSUR)47, 1 (2014), 1–44
work page 2014
-
[8]
Chee-Kheng Ch’ng, Chee Seng Chan, and Cheng-Lin Liu. 2020. Total-text: toward orientation robustness in scene text detection.International Journal on Document Analysis and Recognition (IJDAR)23, 1 (2020), 31–52
work page 2020
Show all 57 references
-
[9]
Denis Coquenet, Clément Chatelain, and Thierry Paquet. 2023. DAN: a segmentation-free document attention network for handwritten document recog- nition.IEEE Transactions on Pattern Analysis and Machine Intelligence(2023)
2023
-
[10]
William Morris Davis. 1893. The topographic maps of the United States Geologi- cal Survey.Science534 (1893), 225–227
-
[11]
Ankush Gupta, Andrea Vedaldi, and Andrew Zisserman. 2016. Synthetic data for text localisation in natural images. InProceedings of the IEEE conference on computer vision and pattern recognition. 2315–2324
2016
-
[12]
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. 2017. Mask r-cnn. InProceedings of the IEEE international conference on computer vision. 2961–2969
2017
-
[13]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. InProceedings of the IEEE conference on computer vision and pattern recognition. 770–778
2016
-
[14]
2011.Map of a nation: A biography of the Ordnance Survey
Rachel Hewitt. 2011.Map of a nation: A biography of the Ordnance Survey. Granta Publications
2011
-
[15]
Mingxin Huang, Yuliang Liu, Zhenghao Peng, Chongyu Liu, Dahua Lin, Sheng- gao Zhu, Nicholas Yuan, Kai Ding, and Lianwen Jin. 2022. SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recog- nition. InProceedings of the IEEE/CVF Conference on...
2022
-
[16]
Mingxin Huang, Jiaxin Zhang, Dezhi Peng, Hao Lu, Can Huang, Yuliang Liu, Xiang Bai, and Lianwen Jin. 2023. Estextspotter: Towards better scene text spotting with explicit synergy in transformer. InProceedings of the IEEE/CVF International Conference on Computer Vision. 19495–19505
2023
-
[17]
Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and CV Jawahar. 2019. Icdar2019 competition on scanned receipt ocr and information extraction. In2019 International Conference on Document Analysis and Recognition (ICDAR). IEEE, 1516–1520
2019
-
[18]
Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014. Synthetic data and artificial neural networks for natural scene text recognition. arXiv preprint arXiv:1406.2227(2014)
2014 arXiv
-
[19]
Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ra- maseshan Chandrasekhar, Shijian Lu, et al. 2015. ICDAR 2015 competition on robust reading. In2015 13th international conference on...
2015
-
[20]
Taeho Kil, Seonghyeon Kim, Sukmin Seo, Yoonsik Kim, and Daehee Kim. 2023. Towards unified scene text spotting based on sequence generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 15223– 15232
2023
-
[21]
Jina Kim, Zekun Li, Yijun Lin, Min Namgung, Leeje Jang, and Yao-Yi Chiang
-
[22]
Seonghyeon Kim, Seung Shin, Yoonsik Kim, Han-Cheol Cho, Taeho Kil, Jae- heung Surh, Seunghyun Park, Bado Lee, and Youngmin Baek. 2022. DEER: Detection-agnostic End-to-End Recognizer for Scene Text Spotting.arXiv preprint arXiv:2203.05122(2022)
2022 arXiv
-
[23]
Yair Kittenplon, Inbal Lavi, Sharon Fogel, Yarin Bar, R Manmatha, and Pietro Perona. 2022. Towards weakly-supervised text spotting using a multi-task transformer. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4604–4613
2022
-
[24]
Harold W Kuhn. 1955. The Hungarian method for the assignment problem. Naval research logistics quarterly2, 1-2 (1955), 83–97
1955
-
[25]
Robert A Lamb. 1961. The Sanborn map: a tool for the geographer. (1961)
1961
-
[26]
2017.QGIS python programming cookbook
Joel Lawhead. 2017.QGIS python programming cookbook. Packt Publishing Ltd
2017
-
[27]
Huali Li, Jun Liu, and Xiran Zhou. 2018. Intelligent map reader: A framework for topographic map understanding with deep learning and gazetteer.IEEE Access6 (2018), 25363–25376
2018
-
[28]
Minghao Li, Tengchao Lv, Jingye Chen, Lei Cui, Yijuan Lu, Dinei Florencio, Cha Zhang, Zhoujun Li, and Furu Wei. 2023. Trocr: Transformer-based optical char- acter recognition with pre-trained models. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 13094–13102
2023
-
[29]
Zekun Li, Yao-Yi Chiang, Sasan Tavakkol, Basel Shbita, Johannes H Uhl, Stefan Leyk, and Craig A Knoblock. 2020. An automatic approach for generating rich, linked geo-metadata from historical map images. InProceedings of the 26th ACM SIGKDD International Conference on Knowledge...
2020
-
[30]
Zekun Li, Runyu Guan, Qianmu Yu, Yao-Yi Chiang, and Craig A Knoblock. 2021. Synthetic Map Generation to Provide Unlimited Training Data for Historical Map Text Detection. InProceedings of the 4th ACM SIGSPATIAL International Workshop on AI for Geographic Knowledge Discovery. 17–26
2021
-
[31]
Minghui Liao, Guan Pang, Jing Huang, Tal Hassner, and Xiang Bai. 2020. Mask textspotter v3: Segmentation proposal network for robust scene text spotting. InComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XI 16. Springer, 706–722
2020
-
[32]
Xuebo Liu, Ding Liang, Shi Yan, Dagui Chen, Yu Qiao, and Junjie Yan. 2018. Fots: Fast oriented text spotting with a unified network. InProceedings of the IEEE conference on computer vision and pattern recognition. 5676–5685
2018
-
[33]
Yuliang Liu, Hao Chen, Chunhua Shen, Tong He, Lianwen Jin, and Liangwei Wang. 2020. Abcnet: Real-time scene text spotting with adaptive bezier-curve network. Inproceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9809–9818
2020
-
[34]
Yuliang Liu, Chunhua Shen, Lianwen Jin, Tong He, Peng Chen, Chongyu Liu, and Hao Chen. 2021. Abcnet v2: Adaptive bezier-curve network for real-time end-to-end text spotting.arXiv preprint arXiv:2105.03620(2021)
2021 arXiv
-
[35]
Francesco Lombardi and Simone Marinai. 2020. Deep learning for historical document analysis and recognition—A survey.Journal of Imaging6, 10 (2020), 110
2020
-
[36]
Nibal Nayef, Fei Yin, Imen Bizid, Hyunsoo Choi, Yuan Feng, Dimosthenis Karatzas, Zhenbo Luo, Umapada Pal, Christophe Rigaud, Joseph Chazalon, et al
-
[37]
Rhett M Olson, Jina Kim, and Yao-Yi Chiang. 2023. An Automatic Approach to Finding Geographic Name Changes on Historical Maps. InProceedings of the 31st ACM International Conference on Advances in Geographic Information Systems. 1–2
2023
-
[38]
Liang Qiao, Sanli Tang, Zhanzhan Cheng, Yunlu Xu, Yi Niu, Shiliang Pu, and Fei Wu. 2020. Text perceptron: Towards end-to-end arbitrary-shaped text spotting. In Proceedings of the AAAI conference on artificial intelligence, Vol. 34. 11899–11907
2020
-
[39]
Archan Ray, Ziwen Chen, Ben Gafford, Nathan Gifford, J Jai Kumar, Abyaya Lamsal, Liam Niehus-Staab, Jerod Weinman, and Erik Learned-Miller. 2018. His- torical map annotations for text detection and recognition.Grinnell College, Grinnell, Iowa, Tech. Rep(2018)
2018
-
[40]
Roi Ronen, Shahar Tsiper, Oron Anschel, Inbal Lavi, Amir Markovitz, and R Manmatha. 2022. Glass: Global to local attention for scene-text spotting. In Computer Vision–ECCV 2022: 17th European Conference, Tel A viv, Israel, October 23–27, 2022, Proceedings, Part XXVIII. Springe...
2022
-
[41]
2002.David Rumsey map collection
David Rumsey. 2002.David Rumsey map collection. Cartography Associates
2002
-
[42]
David Rumsey. 2008. David Rumsey historical map collection
2008
-
[43]
Joan Andreu Sánchez, Verónica Romero, Alejandro H Toselli, Mauricio Villegas, and Enrique Vidal. 2019. A set of benchmarks for handwritten text recognition on historical documents.Pattern Recognition94 (2019), 122–134
2019
-
[44]
Baoguang Shi, Xiang Bai, and Cong Yao. 2016. An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition.IEEE transactions on pattern analysis and machine intelligence39, 11 (2016), 2298–2304
2016
-
[45]
Takahiro Shima, Kengo Terasawa, and Toshio Kawashima. 2011. Image Process- ing for Historical Newspaper Archives. InProceedings of the 2011 Workshop on Historical Document Imaging and Processing (HIP ’11). Association for Computing Machinery, New York, NY, USA, 127–132
2011
-
[46]
Wei Ren Tan, Chee Seng Chan, Hernan Aguirre, and Kiyoshi Tanaka. 2019. Improved ArtGAN for Conditional Synthesis of Natural Image and Artwork. IEEE Transactions on Image Processing28, 1 (2019), 394–409. https://doi.org/10. 1109/TIP.2018.2866698
2019
-
[47]
Hao Wang, Pu Lu, Hui Zhang, Mingkun Yang, Xiang Bai, Yongchao Xu, Mengchao He, Yongpan Wang, and Wenyu Liu. 2020. All you need is boundary: Toward arbitrary-shaped text spotting. InProceedings of the AAAI conference on artificial intelligence, Vol. 34. 12160–12167
2020
-
[48]
Pengfei Wang, Chengquan Zhang, Fei Qi, Shanshan Liu, Xiaoqiang Zhang, Pengyuan Lyu, Junyu Han, Jingtuo Liu, Errui Ding, and Guangming Shi. 2021. Pgnet: Real-time arbitrarily-shaped text spotting with point gathering network. InProceedings of the AAAI Conference on Artificial I...
2021
-
[49]
Tao Wang, David J Wu, Adam Coates, and Andrew Y Ng. 2012. End-to-end text recognition with convolutional neural networks. InProceedings of the 21st international conference on pattern recognition (ICPR2012). IEEE, 3304–3308
2012
-
[50]
Jerod Weinman, Ziwen Chen, Ben Gafford, Nathan Gifford, Abyaya Lamsal, and Liam Niehus-Staab. 2019. Deep neural networks for text detection and recognition in historical maps. In2019 International Conference on Document Analysis and Recognition (ICDAR). IEEE, 902–909
2019
-
[51]
Linjie Xing, Zhi Tian, Weilin Huang, and Matthew R Scott. 2019. Convolutional character networks. InProceedings of the IEEE/CVF international conference on computer vision. 9126–9136
2019
-
[52]
Maoyuan Ye, Jing Zhang, Shanshan Zhao, Juhua Liu, Tongliang Liu, Bo Du, and Dacheng Tao. 2022. DeepSolo: Let Transformer Decoder with Explicit Points Solo for Text Spotting.arXiv preprint arXiv:2211.10772(2022)
2022 arXiv
-
[53]
Maoyuan Ye, Jing Zhang, Shanshan Zhao, Juhua Liu, Tongliang Liu, Bo Du, and Dacheng Tao. 2023. Deepsolo: Let transformer decoder with explicit points solo for text spotting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 19348–19357. Hyper...
2023
-
[54]
Xiang Zhang, Yongwen Su, Subarna Tripathi, and Zhuowen Tu. 2022. Text Spotting Transformers. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 9519–9528
2022
-
[55]
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2020. Deformable detr: Deformable transformers for end-to-end object detection.arXiv preprint arXiv:2010.04159(2020). A Experimental Settings We adopt ResNet-50 [13] as the backbone to extract multi-scale ...
2020 arXiv
-
[2017]
In2017 14th IAPR international conference on document analysis and recognition (ICDAR), Vol
Icdar2017 robust reading challenge on multi-lingual scene text detection and script identification-rrc-mlt. In2017 14th IAPR international conference on document analysis and recognition (ICDAR), Vol. 1. IEEE, 1454–1459
-
[2023]
The mapKurator System: A Complete Pipeline for Extracting and Linking Text from Historical Maps.arXiv preprint arXiv:2306.17059(2023)
2023 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.