REVIEW 3 major objections 5 minor 53 references
Distance transform regression for spatially-aware deep semantic segmentation
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The paper claims that adding signed-distance-transform regression to a fully convolutional network's segmentation loss makes its predictions smoother and more accurate, at almost no architectural cost.
desk verdict A simple and plausible auxiliary loss that improves segmentation boundaries, but the paper's 'significant improvements' claim outruns its statistics—most gains sit within one standard deviation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The signed distance transform (SDT) of a binary mask assigns each foreground pixel its Euclidean distance to the nearest background pixel and each background pixel the negative of that distance; here the SDT is computed per class, clipped to avoid out-of-receptive-field dependencies, and normalized to $[-1,1]$ by a hardtanh. The carrying mechanism is the cascaded multi-task head: the network's last feature maps are used first to regress all SDTs under an L1 loss, then those distance maps are concatenated with the feature maps and passed through one convolution and a softmax for the classification loss. This forces the network to represent, for every pixel, not only 'which class' but 'how far from every class boundary', and the concatenation lets the classifier use the distances as an auxiliary spatial feature. The paper shows that the product of this machinery is smoother boundaries, fewer holes in closed shapes, and reduced classification noise, with a hyperparameter $\lambda$ balancing the two losses.
What would settle it
Retrain the same models on the same splits several times with and without the SDT loss, keeping all hyperparameters fixed, and test whether the paired overall-accuracy differences are consistently positive and exceed the reported standard deviations; the claim fails if the gain does not reproduce on Vaihingen. A sharper ablation is to replace the SDT target with a constant-valued target in the same multi-task head—if accuracy still improves, the distance signal is not what drives the regularization.
Extended reading notes
Core claim
On its own terms, the paper's claim is that training a segmentation network to approximate the signed distance transform of its labels back-propagates spatial cues that implicitly regularize the predicted segmentation. The proposed network head outputs per-class distance maps from the last feature maps, passes them through a hardtanh to keep values in $[-1,1]$, and concatenates these distance predictions with the same feature maps before a final convolution and softmax. The total loss is $L = \mathrm{NLLLoss}(Z_{\mathrm{seg}},Y_{\mathrm{seg}}) + \lambda L_1(Z_{\mathrm{dist}},Y_{\mathrm{dist}})$, where the first term is the usual pixel-wise classification loss. The paper reports quantitative gains such as Vaihingen overall accuracy going from $90.11 \pm 0.11$ to $90.31 \pm 0.12$, CamVid mean IoU rising by up to 2.1 points on DenseUNet, and large IoU gains on the INRIA building benchmark, alongside qualitative improvements in boundary sharpness and shape connectivity. It also notes that regression on the distance transform alone performs worse than classification, and that regressing the binary masks directly does not improve segmentation, so the SDT's spatial content is the active ingredient. On the RGB-D SUN RGB-D set, overall accuracy and average precision improve while average IoU slightly falls.
Load-bearing premise
The central claim of consistent improvement rests on the assumption that score differences on the order of a tenth of a point (e.g., Vaihingen 90.11±0.11 vs 90.31±0.12) and slightly negative changes on one metric (SUN RGB-D AIoU 39.0→38.9) are statistically meaningful rather than run-to-run noise, since the paper reports no significance tests.
Editorial extensions
If this is right
- Adding SDT regression raises reported overall accuracy and per-class F1 on ISPRS Vaihingen and Potsdam (Vaihingen OA 90.11 to 90.31; Potsdam OA 91.85 to 92.22).
- On CamVid, mean IoU improves by 0.5, 1.9, and 2.1 points for PSPNet-50, PSPNet-101, and DenseUNet respectively, while accuracy improves by 0.2, 0.7, and 1.0 points.
- On INRIA aerial building labeling, per-city IoU improves by several points relative to the same SegNet baseline, with cleaner building shapes and fewer false-positive buildings.
- On SUN RGB-D, the FuseNet model's overall accuracy (76.8 to 77.0) and average precision (55.3 to 56.5) improve, while average IoU goes from 39.0 to 38.9.
- The recipe transfers to arbitrary FCN heads because it only changes the final layers, and the paper's binary-mask ablation indicates the distance signal, not just the extra supervision, is responsible for the gains.
Reading between the lines
- The $\lambda \approx 2$ gradient-balancing observation suggests a transferable rule for auxiliary losses: scale an extra regression task so its gradient norm matches the main task; this could be tested on depth, normal, or contour prediction heads.
- Because the SDT is per-class, the memory and compute overhead grows with the number of classes; on very large label sets the method may need to regress distances for sampled or merged classes, which the paper does not address.
- The paper only evaluates semantic segmentation; the same proximity-to-boundary cue should transfer to panoptic or instance segmentation, where object-level SDTs are available, though the paper does not claim this.
- The negative effect seen on CamVid's road and sky classes hints that noisy or void-heavy distance labels can dilute the benefit; testing on datasets with many thin, elongated classes (e.g., lane markings, sidewalks) would map the method's failure boundary.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a multi-task extension of fully convolutional semantic segmentation networks in which, in addition to the usual per-pixel classification loss, the network regresses a truncated, normalized signed distance transform (SDT) of each class mask; the regressed distances are concatenated with the penultimate feature maps and fed to the final classification layer. The total loss is Eq. (2), L = NLL + λ L1. Experiments across five datasets (ISPRS Vaihingen/Potsdam, INRIA, SUN RGB-D, DFC2015, CamVid) with SegNet, PSPNet, FuseNet, and DenseUNet variants are reported. The paper claims significant and consistent quantitative improvements in segmentation accuracy as well as qualitatively smoother boundaries.
Significance. If the empirical claim held, this would be a useful, low-overhead regularization: it requires no architectural redesign beyond a small branch, adds only a hyper-parameter λ and a clipping threshold, and is architecture-agnostic. The authors include useful ablations (regression-only, mask regression, λ sweep) and state that all baselines receive the same extra convolutional layer, which addresses a common confound. However, the central quantitative claim is currently supported only by small metric deltas without error bars or significance tests, and one headline metric moves in the opposite direction; the evidence as presented is not sufficient to establish 'significant improvements'.
major comments (3)
- [§4.4, Tables 1, 4, 5, 6] The abstract and conclusion assert 'significant improvements' and 'consistent quantitative improvements', but no significance test or paired comparison is reported anywhere, and most tables lack error bars. On Vaihingen, overall accuracy changes from 90.11±0.11 to 90.31±0.12, a delta smaller than the reported standard deviation, and on SUN RGB-D the average IoU falls from 39.0 to 38.9 even though the text says results improve. Please provide repeated-run variance estimates, paired significance tests (e.g., per-tile/per-image deltas) or confidence intervals, and revise the headline claims to match what the evidence supports. This is load-bearing because the paper's central claim is the quantitative improvement.
- [§4.3–§4.4, Fig. 8] The value of λ used in each dataset/experiment is not reported in Tables 1, 2, 4, 5, and 6, and the clipping threshold of the SDT is never specified; Fig. 8 shows a λ sweep only for Vaihingen. Without these hyperparameters and the exact protocol (e.g., whether the same λ was used for all datasets), the experiments cannot be reproduced or compared across rows. Please report the selected λ per experiment and, ideally, release the exact configuration.
- [§4.4, Table 6] The claim of 'consistent moderate improvements on all classes' for the DenseUNet variant is contradicted by Table 6: the fence class drops from 26.5 to 23.7, and several classes also degrade in the PSPNet rows (e.g., road and sky). Please revise the per-class characterization and discuss where the method hurts rather than helps.
minor comments (5)
- [Table 1] The Potsdam block does not report standard deviations while the Vaihingen block does; please state whether Potsdam results are single runs and, if so, acknowledge the lack of variance estimates.
- [Table 4] The header contains 'F useNet*' with an extra space; also, 'average IoU' is abbreviated as AIoU, which should be defined in the caption or text.
- [§5] The last paragraph contains the typo 'oly Jégou' for 'only Jégou'.
- [§3.2] Equation (2) would be clearer with explicit spaces around the comma-separated arguments of the loss terms, e.g., NLLLoss(Z_seg, Y_seg) and L1(Z_dist, Y_dist).
- [Fig. 8] In the manuscript version provided to me, Fig. 8 is not reproduced; please verify the figure is included and that its axes are labeled (λ value vs. relative improvement).
Circularity Check
No circularity: SDT targets are deterministic transforms of the labels and the claimed gains are evaluated on held-out benchmarks.
full rationale
The paper's derivation chain is self-contained. The signed distance transform targets are computed deterministically from the ground-truth label masks via Eq. (1), using a linear-time exact Euclidean distance transform (Maurer et al., 2003); they are not derived from, fitted to, or defined in terms of the network's segmentation output. The multi-task loss in Eq. (2) combines a standard NLL classification loss with an L1 regression loss on these precomputed distance maps, so the regression target is an independent label-derived quantity rather than a renamed prediction. Improvements are assessed against classification-only baselines on held-out test sets (ISPRS Vaihingen/Potsdam, INRIA, SUN RGB-D, DFC2015, CamVid), with the same architectures and an added convolutional layer for comparability, so the reported 'predictions' are not equal by construction to fitted inputs. The self-citation to Audebert et al. (2017) appears only in related work as an example of segment-before-detect pipelines and is not load-bearing for the method or its claims. The paper also explicitly credits Bischke et al. (2017) as an independent test of the same idea and compares against it, rather than importing an unverified uniqueness claim from the authors' own prior work. The only mild methodological concern is hyperparameter tuning of lambda on the Vaihingen dataset (Fig. 8 and Table 1), but that is ordinary hyperparameter selection, not a fitted parameter being relabeled as a prediction, and it does not make the central claim equivalent to its inputs. Statistical-significance concerns about small metric deltas are an evidentiary issue, not a circularity issue. Therefore no circular step is present.
Assumptions & free parameters
free parameters (2)
- SDT loss weight lambda =
Explored on Vaihingen: peaks at 0.5 and 2; lambda=2 reported as robust. Per-dataset final values not explicitly stated.
- SDT clipping threshold =
Not reported
assumptions (4)
- standard math Signed Euclidean distance transform can be computed exactly and in linear time for each class mask (Maurer et al. 2003).
- domain assumption Clipped and normalized SDTs retain enough spatial information to be a useful auxiliary learning signal.
- domain assumption A single scalar lambda can balance the gradient magnitudes of classification and regression losses across datasets.
- domain assumption Adding one convolutional layer after concatenating distance predictions with features is sufficient and fair across baselines.
Cite this review
Pith. "Pith review of Distance transform regression for spatially-aware deep semantic segmentation." pith.science (2026). https://pith.science/paper/AG35C4KB
@misc{pith2026190901671,
author = {Pith},
title = {Pith review of: Distance transform regression for spatially-aware deep semantic segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/AG35C4KB}},
note = {Machine review of arXiv:1909.01671}
}
read the original abstract
Understanding visual scenes relies more and more on dense pixel-wise classification obtained via deep fully convolutional neural networks. However, due to the nature of the networks, predictions often suffer from blurry boundaries and ill-segmented shapes, fueling the need for post-processing. This work introduces a new semantic segmentation regularization based on the regression of a distance transform. After computing the distance transform on the label masks, we train a FCN in a multi-task setting in both discrete and continuous spaces by learning jointly classification and distance regression. This requires almost no modification of the network structure and adds a very low overhead to the training process. Learning to approximate the distance transform back-propagates spatial cues that implicitly regularizes the segmentation. We validate this technique with several architectures on various datasets, and we show significant improvements compared to competitive baselines.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in ":" * " " * FUNCTION f...
-
[2]
PyTorch : Tensors and Dynamic neural networks in Python with strong GPU acceleration, 2016. http://pytorch.org/
work page 2016
-
[3]
Anurag Arnab and Philip H. S. Torr. Pixelwise Instance Segmentation With a Dynamically Instantiated Network . In IEEE Conference on Computer Vision and Pattern Recognition ( CVPR ) , pages 441--450, 2017
work page 2017
-
[4]
Nicolas Audebert, Bertrand Le Saux, and S\'ebastien Lef\`evre. Segment-before- Detect : Vehicle Detection and Classification through Semantic Segmentation of Aerial Images . Remote Sensing, 9 0 (4): 0 368, April 2017. doi:10.3390/rs9040368
-
[5]
SegNet : A Deep Convolutional Encoder - Decoder Architecture for Scene Segmentation
Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. SegNet : A Deep Convolutional Encoder - Decoder Architecture for Scene Segmentation . IEEE Transactions on Pattern Analysis and Machine Intelligence, 39 0 (12): 0 2481--2495, December 2017. ISSN 0162-8828. doi:10.1109/TPAMI.2016.2644615
arXiv 2017
-
[6]
Semantic Segmentation With Boundary Neural Fields
Gedas Bertasius, Jianbo Shi, and Lorenzo Torresani. Semantic Segmentation With Boundary Neural Fields . In IEEE Conference on Computer Vision and Pattern Recognition ( CVPR ) , pages 3602--3610, 2016
work page 2016
-
[7]
Multi-Task Learning for Segmentation of Building Footprints with Deep Neural Networks
Benjamin Bischke, Patrick Helber, Joachim Folz, Damian Borth, and Andreas Dengel. Multi- Task Learning for Segmentation of Building Footprints with Deep Neural Networks . arXiv:1709.05932 [cs], September 2017
work page Pith review arXiv 2017
-
[8]
Brostow, Julien Fauqueur, and Roberto Cipolla
Gabriel J. Brostow, Julien Fauqueur, and Roberto Cipolla. Semantic object classes in video: A high-definition ground truth database. Pattern Recognition Letters, 30 0 (2): 0 88--97, January 2009. ISSN 0167-8655. doi:10.1016/j.patrec.2008.04.005
Show all 53 references
-
[9]
Processing of Extremely High - Resolution LiDAR and RGB Data : Outcome of the 2015 IEEE GRSS Data Fusion Contest Part A : 2- D Contest
Manuel Campos-Taberner , Adriana Romero-Soriano , Carlo Gatta, Gustau Camps-Valls , Adrien Lagrange, Bertrand Le Saux, Anne Beaup\`ere, Alexandre Boulch, Adrien Chan-Hon-Tong , St\'ephane Herbin, Hicham Randrianarivo, Marin Ferecatu, Michal Shimoni, Gabriele Moser, and Devis T...
2015
-
[10]
DCAN : Deep Contour - Aware Networks for Accurate Gland Segmentation
Hao Chen, Xiaojuan Qi, Lequan Yu, and Pheng-Ann Heng. DCAN : Deep Contour - Aware Networks for Accurate Gland Segmentation . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 2487--2496, 2016 a
2016
-
[11]
L. C. Chen, J. T. Barron, G. Papandreou, K. Murphy, and A. L. Yuille. Semantic Image Segmentation with Task - Specific Edge Detection Using CNNs and a Discriminatively Trained Domain Transform . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition ,...
2016 doi
-
[12]
DeepLab : Semantic Image Segmentation with Deep Convolutional Nets , Atrous Convolution , and Fully Connected CRFs
Liang-Chieh Chen, Georges Papandreou, Murphy, Kevin , and Yuille, Alan . DeepLab : Semantic Image Segmentation with Deep Convolutional Nets , Atrous Convolution , and Fully Connected CRFs . IEEE Transactions on Pattern Analysis and Machine Intelligence, 40 0 (4): 0 834--848, A...
2018
-
[13]
Cheng, G
D. Cheng, G. Meng, S. Xiang, and C Pan. FusionNet : Edge Aware Deep Convolutional Networks for Semantic Segmentation of Remote Sensing Harbor Images . IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 10 0 (12): 0 5769--5783, December 2017. ISSN...
2017
-
[14]
The Cityscapes Dataset for Semantic Urban Scene Understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The Cityscapes Dataset for Semantic Urban Scene Understanding . In Proceedings of the 2016 IEEE Conference on Computer Vision and Patter...
2016 doi
-
[15]
ImageNet : A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei . ImageNet : A large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition ( CVPR ) , pages 248--255, June 2009. doi:10.1109/CVPR.2009.5206848
2009
-
[16]
Mark Everingham, S. M. Ali Eslami, Luc Van Gool, Christopher K. I. Williams, John Winn, and Andrew Zisserman. The Pascal Visual Object Classes Challenge : A Retrospective . International Journal of Computer Vision, 111 0 (1): 0 98--136, June 2014. ISSN 0920-5691, 1573-1405. do...
2014 doi
-
[17]
A Review on Deep Learning Techniques Applied to Semantic Segmentation
Alberto Garcia-Garcia , Sergio Orts-Escolano , Sergiu Oprea, Victor Villena-Martinez , and Jose Garcia-Rodriguez . A Review on Deep Learning Techniques Applied to Semantic Segmentation . arXiv:1704.06857 [cs], April 2017
2017 arXiv
-
[18]
Boundary-aware Instance Segmentation
Zeeshan Hayder, Xuming He, and Mathieu Salzmann. Boundary-aware Instance Segmentation . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017
2017
-
[19]
FuseNet : Incorporating Depth into Semantic Segmentation via Fusion - Based CNN Architecture
Caner Hazirbas, Lingni Ma, Csaba Domokos, and Daniel Cremers. FuseNet : Incorporating Depth into Semantic Segmentation via Fusion - Based CNN Architecture . In Computer Vision ACCV 2016 , pages 213--228. Springer, Cham , November 2016. doi:10.1007/978-3-319-54181-5_14
2016 doi
-
[20]
Delving Deep into Rectifiers : Surpassing Human - Level Performance on ImageNet Classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving Deep into Rectifiers : Surpassing Human - Level Performance on ImageNet Classification . In Proceedings of the IEEE International Conference on Computer Vision , pages 1026--1034, December 2015. doi:10.1109/ICCV.2015.123
2015 doi
-
[21]
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition ( CVPR ) , pages 770--778, Las Vegas, United States, June 2016. doi:10.1109/CVPR.2016.90
2016 doi
-
[22]
Mask R - CNN
Kaiming He, Georgia Gkioxari, Piotr Doll\'ar, and Ross Girshick. Mask R - CNN . In Proceedings of the International Conference on Computer Vision , March 2017
2017
-
[23]
Weinberger, and Laurens van der Maaten
Gao Huang, Zhuang Liu, Kilian Q. Weinberger, and Laurens van der Maaten . Densely connected convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , volume 1, page 3, 2017
2017
-
[24]
The One Hundred Layers Tiramisu : Fully Convolutional DenseNets for Semantic Segmentation
Simon J\'egou, Michael Drozdzal, David Vazquez, Adriana Romero, and Yoshua Bengio. The One Hundred Layers Tiramisu : Fully Convolutional DenseNets for Semantic Segmentation . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops ( CVPRW ) ,...
2017 doi
-
[25]
Incorporating Depth into both CNN and CRF for Indoor Semantic Segmentation
Jindong Jiang, Zhijun Zhang, Yongqian Huang, and Lunan Zheng. Incorporating Depth into both CNN and CRF for Indoor Semantic Segmentation . arXiv:1705.07383 [cs], May 2017
2017 arXiv
-
[26]
SciPy : Open source scientific tools for Python
Eric Jones, Travis Oliphant, Pearu Peterson, and others . SciPy : Open source scientific tools for Python . http://www.scipy.org/, 2001
2001
-
[27]
Pushing the Boundaries of Boundary Detection using Deep Learning
Iasonas Kokkinos. Pushing the Boundaries of Boundary Detection using Deep Learning . arXiv:1511.07386 [cs], November 2015
2015 arXiv
-
[28]
Hoang Ngan Le, Kha Gia Quach, Khoa Luu, Chi Nhan Duong, and Marios Savvides
TT. Hoang Ngan Le, Kha Gia Quach, Khoa Luu, Chi Nhan Duong, and Marios Savvides. Reformulating Level Sets as Deep Recurrent Neural Network Approach to Semantic Segmentation . IEEE Transactions on Image Processing, 27 0 (5): 0 2393--2407, May 2018. ISSN 1057-7149. doi:10.1109/T...
2018
-
[29]
Efficient piecewise training of deep structured models for semantic segmentation
Guosheng Lin, Chunhua Shen, Anton Van Den Hengel, and Ian Reid. Efficient piecewise training of deep structured models for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition ( CVPR ) , pages 3194--3203, Las Vegas, United Sta...
2016 doi
-
[30]
Lawrence Zitnick
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\'ar, and C. Lawrence Zitnick. Microsoft COCO : Common Objects in Context . In David Fleet, Tomas Pajdla, Bernt Schiele, and Tinne Tuytelaars, editors, Computer Vision ECCV 2014 , n...
2014 doi
-
[31]
Richer Convolutional Features for Edge Detection
Yun Liu, Ming-Ming Cheng, Xiaowei Hu, Kai Wang, and Xiang Bai. Richer Convolutional Features for Edge Detection . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 3000--3009, 2017
2017
-
[32]
Deep Learning Markov Random Field for Semantic Segmentation
Ziwei Liu, Xiaoxiao Li, Ping Luo, Chen Change Loy, and Xiaoou Tang. Deep Learning Markov Random Field for Semantic Segmentation . IEEE Transactions on Pattern Analysis and Machine Intelligence, 40 0 (8): 0 1814--1828, August 2018. ISSN 0162-8828. doi:10.1109/TPAMI.2017.2737535
2018
-
[33]
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition ( CVPR ) , pages 3431--3440, June 2015. doi:10.1109/CVPR.2015.7298965
2015
-
[34]
Can Semantic Labeling Methods Generalize to Any City ? The Inria Aerial Image Labeling Benchmark
Emmanuel Maggiori, Yuliya Tarabalka, Guillaume Charpiat, and Pierre Alliez. Can Semantic Labeling Methods Generalize to Any City ? The Inria Aerial Image Labeling Benchmark . In Proceedings of the IEEE International Symposium on Geoscience and Remote Sensing ( IGARSS ) , July ...
2017
-
[35]
K. K. Maninis, J. Pont-Tuset , P. Arbelaez, and L. Van Gool. Convolutional Oriented Boundaries : From Image Segmentation to High - Level Tasks . IEEE Transactions on Pattern Analysis and Machine Intelligence, 40 0 (4): 0 819--833, April 2018. ISSN 0162-8828. doi:10.1109/TPAMI....
2018
-
[36]
Classification With an Edge : Improving Semantic Image Segmentation with Boundary Detection
Dimitrios Marmanis, Konrad Schindler, Jan Dirk Wegner, Silvano Galliani, Mihai Datcu, and Uwe Stilla. Classification With an Edge : Improving Semantic Image Segmentation with Boundary Detection . ISPRS Journal of Photogrammetry and Remote Sensing, 2017. doi:10.1016/j.isprsjprs...
2017 doi
-
[37]
Maurer, Rensheng Qi, and Vijay Raghavan
Calvin R. Maurer, Rensheng Qi, and Vijay Raghavan. A linear time algorithm for computing exact Euclidean distance transforms of binary images in arbitrary dimensions. IEEE Transactions on Pattern Analysis and Machine Intelligence, 25 0 (2): 0 265--270, February 2003. ISSN 0162...
2003 arXiv
-
[38]
Pinheiro, Tsung-Yi Lin, Ronan Collobert, and Piotr Doll\'ar
Pedro O. Pinheiro, Tsung-Yi Lin, Ronan Collobert, and Piotr Doll\'ar. Learning to Refine Object Segments . In Computer Vision ECCV 2016 , Lecture Notes in Computer Science, pages 75--91. Springer, Cham , October 2016. ISBN 978-3-319-46447-3 978-3-319-46448-0. doi:10.1007/978-3...
2016 doi
-
[39]
3D Graph Neural Networks for RGBD Semantic Segmentation
Xiaojuan Qi, Renjie Liao, Jiaya Jia, Sanja Fidler, and Raquel Urtasun. 3D Graph Neural Networks for RGBD Semantic Segmentation . In Proceedings of the International Conference on Computer Vision , 2017
2017
-
[40]
U- Net : Convolutional Networks for Biomedical Image Segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- Net : Convolutional Networks for Biomedical Image Segmentation . In Nassir Navab, Joachim Hornegger, William M. Wells, and Alejandro F. Frangi, editors, Medical Image Computing and Computer - Assisted Intervention MICCAI 2...
2015 doi
-
[41]
The ISPRS benchmark on urban object classification and 3D building reconstruction
Franz Rottensteiner, Gunho Sohn, Jaewook Jung, Markus Gerke, Caroline Baillard, Sebastien Benitez, and Uwe Breitkopf. The ISPRS benchmark on urban object classification and 3D building reconstruction. ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information S...
2012
-
[42]
Wei Shen, Xinggang Wang, Yan Wang, Xiang Bai, and Z. Zhang. DeepContour : A deep convolutional feature learned by positive-sharing loss for contour detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 3982--3991, June 2015. doi:10...
2015
-
[43]
Very Deep Convolutional Networks for Large - Scale Image Recognition
Karen Simonyan and Andrew Zisserman. Very Deep Convolutional Networks for Large - Scale Image Recognition . In Proceedings of the International Conference on Learning Representations ( ICLR ) , May 2015
2015
-
[44]
Sommer, K
L. Sommer, K. Nie, A. Schumann, T. Schuchert, and J. Beyerer. Semantic labeling for improved vehicle detection in aerial imagery. In IEEE International Conference on Advanced Video and Signal Based Surveillance ( AVSS ) , pages 1--6, August 2017. doi:10.1109/AVSS.2017.8078510
2017
-
[45]
S. Song, S. P. Lichtenberg, and J. Xiao. SUN RGB - D : A RGB - D scene understanding benchmark suite. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 567--576, June 2015. doi:10.1109/CVPR.2015.7298655
2015
-
[46]
Omer Demirel, Lars Malmstr\
Vladim\'ir Ulman, Martin Maska, Klas E. G. Magnusson, Olaf Ronneberger, Carsten Haubold, Nathalie Harder, Pavel Matula, Petr Matula, David Svoboda, Miroslav Radojevic, Ihor Smal, Karl Rohr, Joakim Jald\'en, Helen M. Blau, Oleh Dzyubachyk, Boudewijn Lelieveldt, Pengdong Xiao, Y...
2017
-
[47]
J. Yang, B. Price, S. Cohen, H. Lee, and M. Yang. Object Contour Detection with a Fully Convolutional Encoder - Decoder Network . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition ( CVPR ) , pages 193--202, June 2016. doi:10.1109/CVPR.2016.28
2016 doi
-
[48]
Q. Z. Ye. The signed Euclidean distance transform and its applications. In [1988 Proceedings ] 9th International Conference on Pattern Recognition , pages 495--499 vol.1, November 1988. doi:10.1109/ICPR.1988.28276
1988
-
[49]
Multi- Scale Context Aggregation by Dilated Convolutions
Fisher Yu and Vladlen Koltun. Multi- Scale Context Aggregation by Dilated Convolutions . In Proceedings of the International Conference on Learning Representations ( ICLR ) , November 2015
2015
-
[50]
CASENet : Deep Category - Aware Semantic Edge Detection
Zhiding Yu, Chen Feng, Ming-Yu Liu, and Srikumar Ramalingam. CASENet : Deep Category - Aware Semantic Edge Detection . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 5964--5973, 2017
2017
-
[51]
F. d A. Zampirolli and L. Filipe. A Fast CUDA - Based Implementation for the Euclidean Distance Transform . In International Conference on High Performance Computing Simulation , pages 815--818, July 2017. doi:10.1109/HPCS.2017.123
2017 doi
-
[52]
Pyramid Scene Parsing Network
Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid Scene Parsing Network . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition ( CVPR ) , pages 2881--2890, Honolulu, United States, July 2017. doi:10.1109/CVPR.2017.660
2017 doi
-
[53]
Shuai Zheng, Sadeep Jayasumana, Bernardino Romera-Paredes , Vibhav Vineet, Zhizhong Su, Dalong Du, Chang Huang, and Philip H. S. Torr. Conditional Random Fields as Recurrent Neural Networks . In Proceedings of the IEEE International Conference on Computer Vision , pages 1529--...
2015 doi
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.