Pith. sign in

REVIEW 3 major objections 5 minor 93 references

Benchmarking the Robustness of Semantic Segmentation Models

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper establishes that, within DeepLabv3+, semantic segmentation models that are more accurate on clean images are usually also more robust to realistic image corruptions, with Xception-71 the most robust backbone and the Dense…

desk verdict A careful, large-scale robustness benchmark for segmentation that produces usable design rules, but its reliance on unvalidated synthetic corruptions and single training runs means the headline trend should be read as a strong empirical observation, not a law. read the letter →

arxiv 1908.05005 v3 pith:Q5QHATKU submitted 2019-08-14 cs.CV eess.IV

classification cs.CVeess.IV
keywords semanticsegmentationrobustnessimagecorruptionsDeepLabv3+atrousconvolutionsDensePredictionCellbenchmarkautonomousdriving
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper measures how semantic segmentation models handle image corruptions that a real camera would produce, using nineteen corruption types across three datasets and almost 400,000 images. It centers the study on DeepLabv3+ with six backbones and five retrained architectural ablations, plus several other segmentation models. The central finding is that, contrary to full-image classification, robustness within DeepLabv3+ generally increases with clean-data performance: the strongest backbone, Xception-71, degrades least under most corruptions. The main exception is the Dense Prediction Cell, a search-optimized module that improves clean accuracy but consistently hurts robustness. The paper also derives design rules: keep atrous convolutions and the atrous spatial pyramid pooling module, and treat the Dense Prediction Cell cautiously in safety-critical settings.

What carries the argument

The load-bearing object is the DeepLabv3+ model family used as a controlled testbed: six network backbones (MobileNet-V2, ResNet-50, ResNet-101, Xception-41, Xception-65, Xception-71), each with five architectural ablations (removing atrous convolutions, removing the atrous spatial pyramid pooling module, replacing it with a Dense Prediction Cell, removing the long-range link, and adding global average pooling), all retrained on clean data. Robustness is measured by Corruption Degradation, the sum of mIoU losses over severity levels divided by the same sum for a reference model, and by relative Corruption Degradation, which subtracts the clean-data loss. These metrics turn a large corruption suite into per-property comparisons, isolating the effect of a single architectural change.

What would settle it

Take a camera with measured sensor noise and lens point-spread function, capture the same Cityscapes-like scenes clean and degraded, and rerun the six backbones and the Dense Prediction Cell ablation. If Xception-71 no longer has the lowest corruption degradation, or if the Dense Prediction Cell variant no longer degrades more than the reference model, the paper's central ranking fails outside its synthetic benchmark.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that robustness to image corruptions in DeepLabv3+ semantic segmentation is governed by two factors: backbone strength and architectural module choice. Across Cityscapes, PASCAL VOC 2012, and ADE20K, corruption degradation generally shrinks as clean mean intersection-over-union grows, so the most accurate backbones are usually the most robust, in contrast to the pattern reported for ImageNet classifiers. Atrous convolutions and the long-range link help against blur, noise, and geometric distortion, and the atrous spatial pyramid pooling module is important for decent overall performance. The Dense Prediction Cell, designed purely to maximize clean-data accuracy, consistently reduces robustness, especially for Xception-71, suggesting that clean-only neural architecture search can overfit to the clean objective.

Load-bearing premise

The entire ranking rests on the assumption that the nineteen synthetic corruptions, including the new intensity-dependent noise and PSF blur, behave like the distortions a real deployed camera produces, and the paper does not validate that correspondence against physical camera images.

Editorial extensions

If this is right

  • Practitioners can usually select the most accurate DeepLabv3+ backbone, Xception-71, without paying a robustness penalty, since it has the lowest corruption degradation on all three datasets.
  • Atrous convolutions should be kept in segmentation backbones: removing them consistently increases degradation under blur, noise, and geometric distortion on Cityscapes.
  • The atrous spatial pyramid pooling module is structurally important for robustness, not just accuracy; removing it raises corruption degradation across datasets and backbones.
  • Deploying a Dense Prediction Cell in safety-critical settings is risky: it wins on clean mIoU for Xception-71 but raises corruption degradation to roughly 109 to 115 percent for noise and similar levels for other corruptions.
  • Because clean accuracy and robustness are positively correlated within this architecture, robustness does not have to be traded against performance when selecting segmentation backbones.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension is to feed the same corruption suite to other neural-architecture-search-derived modules: if the Dense Prediction Cell result generalizes, clean-only search objectives should be expected to systematically overfit to clean images, and robustness should become a search objective.
  • The paper's shape-bias speculation invites a direct experiment: measure whether Xception backbones classify corrupted Cityscapes objects with more shape reliance than ResNet backbones, and whether that predicts their robustness advantage.
  • The new intensity-dependent noise model and PSF blur are generated synthetically, so an obvious next check is whether the same rankings survive on images from a real camera whose sensor noise and lens aberrations were measured, since the benchmark does not validate its corruptions against physical captures.
  • If the positive performance-robustness relationship holds for other segmentation architectures, model selection for autonomous driving could use clean mIoU as a rough robustness proxy, reducing the need for exhaustive corruption testing.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents a large-scale robustness benchmark for semantic segmentation, centered on DeepLabv3+ with six network backbones and five architectural ablations, evaluated on Cityscapes, PASCAL VOC 2012, and ADE20K. The corruption suite combines 15 ImageNet-C corruptions with three newly proposed degradations (intensity-dependent camera noise, spatially varying PSF blur, and geometric distortion), yielding almost 400,000 corrupted validation images. Robustness is measured using Corruption Degradation (CD) and relative Corruption Degradation (rCD). The two central claims are that, contrary to full-image classification, DeepLabv3+ robustness often increases with clean-data performance (with Xception-71 the most robust backbone), and that specific architectural properties matter: atrous convolutions and the long-range link generally help, while the Dense Prediction Cell hurts robustness despite improving clean accuracy.

Significance. If the claims hold, this is one of the first and most extensive robustness studies for semantic segmentation, and it provides actionable architectural guidance for deploying segmentation models in safety-critical settings. The paper's strengths are its scale (102 retrained models, three datasets, 19 corruptions), its consistent evaluation protocol, the explicit separation of CD and rCD, and the qualitative and quantitative ablation evidence. The proposed realistic noise and PSF blur models are a step beyond standard Gaussian-only corruptions. However, the external validity of the benchmark depends crucially on whether the synthetic corruptions faithfully represent real camera degradations, and the statistical support for some design rules is currently weak because each configuration is trained only once.

major comments (3)
  1. [Section 3.2 (Eq. 3); Section 3.1] The proposed 'more realistic' corruptions are not validated against real camera data, although the introduction and abstract motivate the benchmark with practical applications such as autonomous driving. The intensity-dependent noise model in Eq. (3) has a free severity parameter w_s but no fit to measured sensor noise (e.g., dark-frame or flat-field statistics), the PSF kernels are generated with Zemax without comparison to a real lens PSF, and the severity calibration is provided only for the noise category (Table A.1); blur, weather, digital, and geometric severities have no SNR or perceptual anchor. Since the headline trend and the design rules (atrous helps, DPC hurts) are averages over these synthetic corruptions, the transfer to deployment is an unsubstantiated leap. I ask for a quantitative validation of the new corruption models against real camera outputs, or for the conclusions to be explicitly restricted to the synthetic corruption suite.
  2. [Section 5.3 (Table 2)] The text states that the most distinct 'statistically significant' results are discussed, but no significance test is reported anywhere and each architecture/ablation is trained only once. Table 2 reports a standard deviation for image noise of 0.2 or less, yet no test statistic, confidence interval, or number of repeated runs is given. Several conclusions rest on small CD differences, especially on ADE20K where the mean CD for most ablations is within 1-2% of 100 (Tables B.9-B.10). The authors should either provide repeated training runs with a paired test over corruptions or rephrase the claim as 'consistent differences' and remove the term 'statistically significant'; otherwise the proposed design rules may be partly driven by optimization noise.
  3. [Section 5.2 (Fig. 4)] The central claim that robustness increases with model performance contrasts the DeepLabv3+ backbone results with results for full-image classification, but the comparison mixes architecture family and task. For non-DeepLab segmentation models in Fig. 4(d), the CD decreases with clean performance while the rCD stays above 100% - the same pattern the authors attribute to classification. The 'most cases' wording is honest, but the claim should identify explicitly that the positive performance-robustness correlation is an intra-architecture finding for DeepLabv3+ with MobileNet-V2 as reference, not a general property of semantic segmentation models. Please make this scope explicit in the abstract and conclusions.
minor comments (5)
  1. [Eqs. (1)-(2)] Equations (1) and (2) sum over s=1 to 5, while the text and table captions state that only the first three severity levels are used for the noise category; please define an explicit severity set in the equation or add a sentence explaining the convention.
  2. [Table 1 / Section 3.2] PSF blur is described as having three severity levels and geometric distortion has no severity levels, but Table 1 reports an 'average mIoU' for each corruption without explaining how the averaging over severity is performed for these two corruptions; please clarify.
  3. [Table 2 caption] The statement 'The standard deviation for image noise is 0.2 or less' is ambiguous: it is unclear whether this is a deviation across images, across severity levels, or across training runs. Please state the source and computation explicitly.
  4. [Fig. 4] The panels label the y-axis 'Corruption Degradation [%]' but plot both CD and rCD, and each legend repeats the reference model name; please split or relabel the panels to avoid confusion.
  5. [Section 5.3] The notation w\o, w/, and w\DPC is visually error-prone; using standard text such as 'without AC' and 'with DPC' would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the robustness benchmark reports measured mIoU/CD/rCD values, and the headline performance-robustness trend is an empirical observation, not a consequence of the definitions or of any self-citation.

full rationale

This paper is an empirical benchmark study. The central claims—that DeepLabv3+ robustness often increases with clean mIoU, that atrous convolutions help against blur and noise, and that DPC reduces robustness—are read off from measured mIoU tables (Tables 1 and 2) and from Corruption Degradation (CD) / relative Corruption Degradation (rCD) values computed via Eqs. (1) and (2). CD and rCD are ratio metrics that normalize a model's degradation by a reference model's degradation, but nothing in these definitions forces the observed ordering of backbones or the ablation effects; the correlation with clean performance is an empirical result about the trained models, not an identity. The synthetic corruption suite (ImageNet-C plus the proposed intensity-dependent noise, PSF blur, and geometric distortion) is a benchmark construction, and its realism relative to real cameras is a legitimate external-validity concern, but it is not a circularity: the results are measured performance under that suite, not quantities fitted to reproduce the conclusion. No parameter is fitted to a subset of the robustness data and then renamed a prediction, no uniqueness theorem is imported from the authors' prior work, and the references contain no load-bearing self-citation chain. The paper is self-contained as a benchmark evaluation, so no circular step is present.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The central claims rest on the representativeness of the corruption protocol, the adequacy of mIoU, CD, and rCD as robustness measures, and the comparability of single training runs. The intensity-noise strength is the only hand-set numeric knob that could materially shift aggregate rankings, and its values are not reported numerically.

free parameters (1)
  • Intensity-dependent noise strength w_s = One value per severity level (1-5), not disclosed numerically
    Equation 3 in Supplement A.2 controls how strongly luminance and chrominance noise are added. It is set by hand and is not calibrated against real camera noise statistics; the paper does not provide sensitivity analysis, so aggregate robustness comparisons depend on this arbitrary severity scaling.
assumptions (3)
  • domain assumption mIoU, CD, and rCD as defined in Equations 1 and 2 are adequate measures of robustness for ranking models.
    All conclusions are expressed through these metrics. The paper chooses one reference model, MobileNet-V2, equal severity averaging, and includes only the first three noise severity levels in the averages.
  • domain assumption The ImageNet-C corruptions and the proposed intensity noise, PSF blur, and geometric distortion represent realistic input degradation.
    Section 3.2 asserts realism but does not validate the new corruption models against real sensor or camera images. The ImageNet-C corruptions were originally designed for full-image classification.
  • ad hoc to paper A single training run per configuration is sufficient to support statements about statistically significant differences in robustness.
    Sections 5.1 and B.2 describe one training per configuration. No random seeds or significance tests are reported, yet Section 5.3 uses the term 'statistically significant'.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Benchmarking the Robustness of Semantic Segmentation Models." pith.science (2026). https://pith.science/paper/Q5QHATKU

@misc{pith2026190805005,
  author       = {Pith},
  title        = {Pith review of: Benchmarking the Robustness of Semantic Segmentation Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q5QHATKU}},
  note         = {Machine review of arXiv:1908.05005}
}
read the original abstract

When designing a semantic segmentation module for a practical application, such as autonomous driving, it is crucial to understand the robustness of the module with respect to a wide range of image corruptions. While there are recent robustness studies for full-image classification, we are the first to present an exhaustive study for semantic segmentation, based on the state-of-the-art model DeepLabv3+. To increase the realism of our study, we utilize almost 400,000 images generated from Cityscapes, PASCAL VOC 2012, and ADE20K. Based on the benchmark study, we gain several new insights. Firstly, contrary to full-image classification, model robustness increases with model performance, in most cases. Secondly, some architecture properties affect robustness significantly, such as a Dense Prediction Cell, which was designed to maximize performance on clean data only.

Figures

Figures reproduced from arXiv: 1908.05005 by the authors.

Figure 1
Figure 1. Results of our ablation study. Here we train the state-of-the-art semantic segmentation model DeepLabv3 [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. A crop of a validation image from Cityscapes corrupted [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Building blocks of DeepLabv3+. Input images are firstly processed by a network backbone, containing atrous convo￾lutions. The backbone output is further processed by a multi-scale processing module (ASPP or DPC). A long-range link concate￾nates early features of the network backbone with encoder output. Finally, the decoder outputs estimates of semantic labels. Our ref￾erence model is shown by regular arrows (i.e., … view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: (a−c) CD and rCD for several network backbones of the DeepLabv3+ architecture evaluated on PASCAL VOC 2012, the Cityscapes dataset, and ADE20K. MobileNet-V2 is the reference model in each case. rCD and CD values below 100 % represent higher robustness than the referenc…
Figure 5
Figure 5. Figure 5: CD evaluated on Cityscapes for the proposed ablated variants of the DeepLabv3 [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

93 extracted references · 64 canonical work pages

  1. [1]

    Mur- ray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete War- den, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng

    Martin Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghe- mawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Mur- ray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete War- den, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. Tensor- Flow: A system f...

  2. [2]

    Anurag Arnab, Ondrej Miksik, and Philip H. S. Torr. On the Robustness of Semantic Segmentation Models to Adversar- ial Attacks. In CVPR, 2018

  3. [3]

    Why do deep convolutional networks generalize so poorly to small image transforma- tions? CoRR, abs/1805.12177, 2018

    Aharon Azulay and Yair Weiss. Why do deep convolutional networks generalize so poorly to small image transforma- tions? CoRR, abs/1805.12177, 2018

  4. [4]

    SegNet: A Deep Convolutional Encoder-Decoder Architec- ture for Image Segmentation

    Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. SegNet: A Deep Convolutional Encoder-Decoder Architec- ture for Image Segmentation. In PAMI, 2017

  5. [5]

    Learning to Remove Rain in Traf- fic Surveillance by Using Synthetic Data

    Chris H Bahnsen, David Vzquez, Antonio M Lpez, and Thomas B Moeslund. Learning to Remove Rain in Traf- fic Surveillance by Using Synthetic Data. In VISI-GRAPP, 2019

  6. [6]

    CNN-Cert: An Efficient Framework for Certifying Robustness of Convolutional Neural Networks

    Akhilan Boopathy, Tsui-Wei Weng, Pin-Yu Chen, Sijia Liu, and Luca Daniel. CNN-Cert: An Efficient Framework for Certifying Robustness of Convolutional Neural Networks. In AAAI, Jan. 2019

  7. [7]

    DeepCorrect: Correcting DNN models against Image Distortions

    Tejas S. Borkar and Lina J. Karam. DeepCorrect: Correcting DNN models against Image Distortions. arXiv:1705.02406 [cs.CV], 2017

  8. [8]

    Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods

    Nicholas Carlini and David Wagner. Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods. In Proceedings of the 10th ACM Workshop on Artificial In- telligence and Security, AISec ’17, pages 3–14, New York, NY , USA, 2017. ACM

Show all 93 references
  1. [9]

    Nicholas Carlini and David A. Wagner. Towards Evaluating the Robustness of Neural Networks. 2017 IEEE Symposium on Security and Privacy (SP), 2017

  2. [10]

    Collins, Yukun Zhu, George Papandreou, Barret Zoph, Florian Schroff, Hartwig Adam, and Jonathon Shlens

    Liang-Chieh Chen, Maxwell D. Collins, Yukun Zhu, George Papandreou, Barret Zoph, Florian Schroff, Hartwig Adam, and Jonathon Shlens. Searching for Efficient Multi-Scale Ar- chitectures for Dense Image Prediction. In NIPS, 2018

  3. [11]

    Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L. Yuille. Semantic Image Segmen- tation with Deep Convolutional Nets and Fully Connected CRFs. In ICLR, volume abs/1412.7062, 2015

  4. [12]

    Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L. Yuille. DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs. In TPAMI, 2017

  5. [13]

    Rethinking Atrous Convolution for Seman- tic Image Segmentation, 2017

    Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam. Rethinking Atrous Convolution for Seman- tic Image Segmentation, 2017

  6. [14]

    Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation

    Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. In ECCV, 2018

  7. [15]

    Domain adaptive faster r-cnn for object de- tection in the wild

    Yuhua Chen, Wen Li, Christos Sakaridis, Dengxin Dai, and Luc Van Gool. Domain adaptive faster r-cnn for object de- tection in the wild. In CVPR, pages 3339–3348, 2018

  8. [16]

    Xception: Deep Learning with Depthwise Separable Convolutions

    Francois Chollet. Xception: Deep Learning with Depthwise Separable Convolutions. In CVPR, 2017

  9. [17]

    Parseval Networks: Improv- ing Robustness to Adversarial Examples

    Moustapha Cisse, Piotr Bojanowski, Edouard Grave, Yann Dauphin, and Nicolas Usunier. Parseval Networks: Improv- ing Robustness to Adversarial Examples. In Doina Precup and Yee Whye Teh, editors, Proceedings of the 34th Inter- national Conference on Machine Learning , Proceedin...

  10. [18]

    The Cityscapes Dataset for Semantic Urban Scene Understanding

    Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The Cityscapes Dataset for Semantic Urban Scene Understanding. In CVPR, 2016

  11. [19]

    Le.Intriguing Properties of Adversarial Exam- ples

    Ekin Dogus Cubuk, Barret Zoph, Samuel Stern Schoenholz, and Quoc V . Le.Intriguing Properties of Adversarial Exam- ples. 2018

  12. [20]

    Dark model adaptation: Semantic image segmentation from daytime to nighttime

    Dengxin Dai and Luc Van Gool. Dark model adaptation: Semantic image segmentation from daytime to nighttime. In ITSC, pages 3819–3824. IEEE, 2018

  13. [21]

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR, 2009

  14. [22]

    A study and comparison of human and deep learning recognition performance under visual distortions

    Samuel Dodge and Lina Karam. A study and comparison of human and deep learning recognition performance under visual distortions. In 2017 26th international conference on computer communication and networks (ICCCN) , pages 1–

  15. [23]

    Dodge and Lina J

    Samuel F. Dodge and Lina J. Karam. Understanding how im- age quality affects deep neural networks. In Quomex, 2016

  16. [24]

    Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John Winn, and Andrew Zisserman. The Pascal Visual Object Classes (VOC) Challenge. In IJCV, 2010

  17. [25]

    A. W. Fitzgibbon. Simultaneous linear estimation of multi- ple view geometry and lens distortion. In CVPR, volume 1, pages I–I, Dec. 2001

  18. [26]

    Geirhos, P

    R. Geirhos, P. Rubisch, C. Michaelis, M. Bethge, F. A. Wich- mann, and W. Brendel. ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In ICLR, May 2019

  19. [27]

    Medina Temme, Jonas Rauber, Heiko H

    Robert Geirhos, Carlos R. Medina Temme, Jonas Rauber, Heiko H. Schtt, Matthias Bethge, and Felix A. Wichmann. Generalisation in humans and deep neural networks. NIPS, abs/1808.08750, 2018

  20. [28]

    Adversarial Examples Are a Natural Consequence of Test Error in Noise

    Justin Gilmer, Nicolas Ford, Nicholas Carlini, and Ekin Cubuk. Adversarial Examples Are a Natural Consequence of Test Error in Noise. In Kamalika Chaudhuri and Rus- lan Salakhutdinov, editors, Proceedings of the 36th Interna- tional Conference on Machine Learning, volume 97 of...

  21. [29]

    Deep Learning

    Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep Learning. MIT Press, 2016

  22. [30]

    Grauman and T

    K. Grauman and T. Darrell. The Pyramid Match Kernel: Dis- criminative Classification with Sets of Image Features. In ICCV, 2005

  23. [31]

    Towards Deep Neural Net- work Architectures Robust to Adversarial Examples

    Shixiang Gu and Luca Rigazio. Towards Deep Neural Net- work Architectures Robust to Adversarial Examples. NIPS Workshop on Deep Learning and Representation Learning , abs/1412.5068, 2014

  24. [32]

    Hypercolumns for object segmentation and fine- grained localization

    Bharath Hariharan, Pablo Arbelez, Ross Girshick, and Jiten- dra Malik. Hypercolumns for object segmentation and fine- grained localization. In CVPR, pages 447–456, 2015

  25. [33]

    Multiple view ge- ometry in computer vision

    Richard Hartley and Andrew Zisserman. Multiple view ge- ometry in computer vision . Cambridge university press, 2003

  26. [34]

    Hasirlioglu, A

    S. Hasirlioglu, A. Kamann, I. Doric, and T. Brandmeier. Test methodology for rain influence on automotive surround sen- sors. In ITSC, pages 2242–2247, Nov. 2016

  27. [35]

    Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition. In ECCV, 2014

  28. [36]

    Delving Deep into Rectifiers: Surpassing Human-Level Per- formance on ImageNet Classification

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving Deep into Rectifiers: Surpassing Human-Level Per- formance on ImageNet Classification. ICCV, pages 1026– 1034, 2015

  29. [37]

    Deep Residual Learning for Image Recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In CVPR, 2016

  30. [38]

    Radiometric CCD camera calibration and noise estimation

    Glenn E Healey and Raghava Kondepudy. Radiometric CCD camera calibration and noise estimation. PAMI, 16(3):267– 276, 1994

  31. [39]

    Benchmarking Neu- ral Network Robustness to Common Corruptions and Per- turbations

    Dan Hendrycks and Thomas Dietterich. Benchmarking Neu- ral Network Robustness to Common Corruptions and Per- turbations. Proceedings of the International Conference on Learning Representations, 2019

  32. [40]

    Henriques and Andrea Vedaldi

    Joo F. Henriques and Andrea Vedaldi. Warped Convolutions: Efficient Invariance to Spatial Transformations. In ICML, 2017

  33. [41]

    Holschneider, R

    M. Holschneider, R. Kronland-Martinet, J. Morlet, and P. Tchamitchian. A Real-Time Algorithm for Signal Analysis with the Help of the Wavelet Transform. In J.-M. Combes, A. Grossmann, and P. Tchamitchian, editors,Wavelets. Time- Frequency Methods and Phase Space, page 286, 1989

  34. [42]

    Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco An- dreetto, and Hartwig Adam

    Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco An- dreetto, and Hartwig Adam. MobileNets: Efficient Con- volutional Neural Networks for Mobile Vision Applications. CoRR, abs/1704.04861, 2017

  35. [43]

    Weinberger

    Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kil- ian Q. Weinberger. Densely Connected Convolutional Net- works. CVPR, pages 2261–2269, 2017

  36. [44]

    Kwiatkowska, Sen Wang, and Min Wu

    Xiaowei Huang, Marta Z. Kwiatkowska, Sen Wang, and Min Wu. Safety Verification of Deep Neural Networks. In CAV, 2017

  37. [45]

    Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

    Ioffe, Sergey and Szegedy, Christian. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In ICML, 2015

  38. [46]

    Computer Vision for Autonomous Vehicles: Problems, Datasets and State-of-the-Art

    Joel Janai, Fatma Gney, Aseem Behl, and Andreas Geiger. Computer Vision for Autonomous Vehicles: Problems, Datasets and State-of-the-Art. Arxiv, 2017

  39. [47]

    Joshi, R

    N. Joshi, R. Szeliski, and D. J. Kriegman. PSF estimation using sharp edge prediction. InCVPR, pages 1–8, June 2008

  40. [48]

    Kamann, S

    A. Kamann, S. Hasirlioglu, I. Doric, T. Speth, T. Brandmeier, and U. T. Schwarz. Test Methodology for Automotive Sur- round Sensors in Dynamic Driving Situations. In 2017 IEEE 85th Vehicular Technology Conference (VTC Spring), pages 1–6, June 2017

  41. [49]

    Tsung-Wei Ke, Michael Maire, and Stella X. Yu. Multigrid Neural Architectures. In CVPR, pages 4067–4075, 2017

  42. [50]

    Imagenet classification with deep convolutional neural net- works

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural net- works. In Advances in neural information processing sys- tems, pages 1097–1105, 2012

  43. [51]

    Be- yond Bags of Features: Spatial Pyramid Matching for Rec- ognizing Natural Scene Categories

    Svetlana Lazebnik, Cordelia Schmid, and Jean Ponce. Be- yond Bags of Features: Spatial Pyramid Matching for Rec- ognizing Natural Scene Categories. In CVPR, Washington, DC, USA, 2006

  44. [52]

    Yann LeCun, Yoshua Bengio, and Geoffrey E. Hinton. Deep learning. In Nature, 2015

  45. [53]

    Gradient-based learning applied to document recog- nition

    Yann Lecun, Lon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recog- nition. Proceedings of the IEEE, 1998

  46. [54]

    Network in net- work

    Min Lin, Qiang Chen, and Shuicheng Yan. Network in net- work. In ICLR, 2014

  47. [55]

    C. Liu, R. Szeliski, S. Bing Kang, C. L. Zitnick, and W. T. Freeman. Automatic Estimation and Removal of Noise from a Single Image. PAMI, 30(2):299–314, Feb. 2008

  48. [56]

    Fully Convolutional Networks for Semantic Segmentation

    Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully Convolutional Networks for Semantic Segmentation. In CVPR, volume abs/1411.4038, 2015

  49. [57]

    Lukas, J

    J. Lukas, J. Fridrich, and M. Goljan. Digital camera identi- fication from sensor pattern noise. IEEE Transactions on In- formation Forensics and Security, 1(2):205–214, June 2006

  50. [58]

    On Detecting Adversarial Perturbations

    Jan Hendrik Metzen, Tim Genewein, V olker Fischer, and Bastian Bischoff. On Detecting Adversarial Perturbations. In ICLR, 2017

  51. [59]

    Michaelis, B

    C. Michaelis, B. Mitzkus, R. Geirhos, E. Rusak, O. Bring- mann, A. S. Ecker, M. Bethge, and W. Brendel. Bench- marking Robustness in Object Detection: Autonomous Driv- ing when Winter is Coming. In Machine Learning for Autonomous Driving Workshop, NeurIPS 2019 , volume 1907074...

  52. [60]

    Visual Quality Enhancement Of Images Under Adverse Weather Conditions

    Jashojit Mukherjee, K Praveen, and Venugopala Madumbu. Visual Quality Enhancement Of Images Under Adverse Weather Conditions. In ITSC, pages 3059–3066. IEEE, 2018

  53. [61]

    Exploring Generaliza- tion in Deep Learning

    Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro. Exploring Generaliza- tion in Deep Learning. In NIPS, 2017

  54. [62]

    Modeling local and global deformations in Deep Learning: Epitomic convolution, Multiple Instance Learn- ing, and sliding window detection

    George Papandreou, Iasonas Kokkinos, and Pierre-Andr Savalle. Modeling local and global deformations in Deep Learning: Epitomic convolution, Multiple Instance Learn- ing, and sliding window detection. InCVPR, pages 390–399, 2015

  55. [63]

    ENet: A Deep Neural Network Ar- chitecture for Real-Time Semantic Segmentation

    Adam Paszke, Abhishek Chaurasia, Sangpil Kim, and Eu- genio Culurciello. ENet: A Deep Neural Network Ar- chitecture for Real-Time Semantic Segmentation. CoRR, abs/1606.02147, 2016

  56. [64]

    Automatic Dif- ferentiation in PyTorch

    Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic Dif- ferentiation in PyTorch. In NIPS Autodiff Workshop, 2017

  57. [65]

    Efficient neural architecture search via parameter sharing

    Hieu Pham, Melody Y Guan, Barret Zoph, Quoc V Le, and Jeff Dean. Efficient neural architecture search via parameter sharing. ICML, 2018

  58. [66]

    Deformable convolutional net- workscoco detection and segmentation challenge 2017 entry

    Haozhi Qi, Zheng Zhang, Bin Xiao, Han Hu, Bowen Cheng, Yichen Wei, and Jifeng Dai. Deformable convolutional net- workscoco detection and segmentation challenge 2017 entry. In ICCV COCO Challenge Workshop, volume 15, 2017

  59. [67]

    Girshick, and Ali Farhadi

    Joseph Redmon, Santosh Kumar Divvala, Ross B. Girshick, and Ali Farhadi. You Only Look Once: Unified, Real-Time Object Detection. In CVPR, pages 779–788, 2016

  60. [68]

    Se- mantic foggy scene understanding with synthetic data.IJCV, 126(9):973–992, 2018

    Christos Sakaridis, Dengxin Dai, and Luc Van Gool. Se- mantic foggy scene understanding with synthetic data.IJCV, 126(9):973–992, 2018

  61. [69]

    Guided Curriculum Model Adaptation and Uncertainty-Aware Eval- uation for Semantic Nighttime Image Segmentation

    Christos Sakaridis, Dengxin Dai, and Luc Van Gool. Guided Curriculum Model Adaptation and Uncertainty-Aware Eval- uation for Semantic Nighttime Image Segmentation. In ICCV, 2019

  62. [70]

    MobileNetV2: Inverted Residuals and Linear Bottlenecks

    Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zh- moginov, and Liang-Chieh Chen. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In CVPR, 2018

  63. [71]

    Overfeat: Integrated recognition, localization and detection using convolutional networks

    Pierre Sermanet, David Eigen, Xiang Zhang, Michal Math- ieu, Robert Fergus, and Yann Lecun. Overfeat: Integrated recognition, localization and detection using convolutional networks. In ICLR, 2014

  64. [72]

    Meet P. Shah. Semantic Segmenta- tion Architectures Implemented in PyTorch. https://github.com/meetshah1995/pytorch-semseg, 2017

  65. [73]

    Intrinsic parameter cali- bration procedure for a (high-distortion) fish-eye lens cam- era with distortion model and accuracy estimation

    Shishir Shah and JK Aggarwal. Intrinsic parameter cali- bration procedure for a (high-distortion) fish-eye lens cam- era with distortion model and accuracy estimation. Pattern Recognition, 29(11):1775–1788, 1996

  66. [74]

    Very Deep Con- volutional Networks for Large-Scale Image Recognition

    Karen Simonyan and Andrew Zisserman. Very Deep Con- volutional Networks for Large-Scale Image Recognition. In ICLR, 2015

  67. [75]

    Feature Quantization for Defending Against Dis- tortion of Images

    Zhun Sun, Mete Ozay, Yan Zhang, Xing Liu, and Takayuki Okatani. Feature Quantization for Defending Against Dis- tortion of Images. In CVPR, June 2018

  68. [76]

    Going deeper with convolutions

    Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In CVPR, 2015

  69. [77]

    Gated-SCNN: Gated Shape CNNs for Semantic Seg- mentation

    Towaki Takikawa, David Acuna, Varun Jampani, and Sanja Fidler. Gated-SCNN: Gated Shape CNNs for Semantic Seg- mentation. ICCV, 2019

  70. [78]

    Examining the Impact of Blur on Recog- nition by Convolutional Networks

    Igor Vasiljevic, Ayan Chakrabarti, and Gregory Shakhnarovich. Examining the Impact of Blur on Recog- nition by Convolutional Networks. arXiv:1611.05760 [cs.CV], abs/1611.05760, 2016

  71. [79]

    Towards Robust CNN- Based Object Detection through Augmentation with Syn- thetic Rain Variations

    Georg V olk, Mueller Stefan, Alexander von Bernuth, Den- nis Hospach, and Oliver Bringmann. Towards Robust CNN- Based Object Detection through Augmentation with Syn- thetic Rain Variations. In ITSC, 2019

  72. [80]

    Modeling and calibration of automated zoom lenses

    Reg G Willson. Modeling and calibration of automated zoom lenses. In Videometrics III, volume 2350, pages 170–187. International Society for Optics and Photonics, 1994

  73. [81]

    Wider or deeper: Revisiting the resnet model for visual recognition

    Zifeng Wu, Chunhua Shen, and Anton Van Den Hengel. Wider or deeper: Revisiting the resnet model for visual recognition. Pattern Recognition, 90:119–133, 2019

  74. [82]

    Enhancing the Perfor- mance of Convolutional Neural Networks on Quality De- graded Datasets

    Jonghwa Yim and Kyung-Ah Sohn. Enhancing the Perfor- mance of Convolutional Neural Networks on Quality De- graded Datasets. DICTA, 2017

  75. [83]

    Delft University of Technology Delft, 1998

    Ian T Young, Jan J Gerbrands, and Lucas J Van Vliet.Funda- mentals of image processing, volume 841. Delft University of Technology Delft, 1998

  76. [84]

    Multi-Scale Context Aggre- gation by Dilated Convolutions

    Fisher Yu and Vladlen Koltun. Multi-Scale Context Aggre- gation by Dilated Convolutions. In ICLR, 2016

  77. [85]

    ICNet for Real-Time Semantic Segmen- tation on High-Resolution Images

    Hengshuang Zhao, Xiaojuan Qi, Xiaoyong Shen, Jianping Shi, and Jiaya Jia. ICNet for Real-Time Semantic Segmen- tation on High-Resolution Images. In ECCV, 2018

  78. [86]

    Pyramid Scene Parsing Network

    Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid Scene Parsing Network. In CVPR, 2017

  79. [87]

    Good- fellow

    Stephan Zheng, Yang Song, Thomas Leung, and Ian J. Good- fellow. Improving the Robustness of Deep Neural Networks via Stability Training. In CVPR, pages 4480–4488, 2016

  80. [88]

    Scene parsing through ade20k dataset

    Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In CVPR, pages 633–641, 2017

  81. [89]

    Semantic under- standing of scenes through the ade20k dataset

    Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fi- dler, Adela Barriuso, and Antonio Torralba. Semantic under- standing of scenes through the ade20k dataset. IJCV, pages 1–20, 2016

  82. [90]

    On Classifi- cation of Distorted Images with Deep Convolutional Neural Networks

    Yiren Zhou, Sibo Song, and Ngai-Man Cheung. On Classifi- cation of Distorted Images with Deep Convolutional Neural Networks. ICASSP, 2017

  83. [91]

    Barret Zoph and Quoc V . Le. Neural Architecture Search with Reinforcement Learning. 2017

  84. [92]

    jpeg compression

    Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le. Learning transferable architectures for scalable image recognition. In CVPR, pages 8697–8710, 2018. Supplemental Material We provide further information about the utilized image corruptions and the conducted experim...

  85. [93]

    On ADE20K, the mIoU de- creases between 1.2 % (Xception-65) and 7.7 % (ResNet- 50)

    and 12.0 % (ResNet-50). On ADE20K, the mIoU de- creases between 1.2 % (Xception-65) and 7.7 % (ResNet- 50). On Cityscapes, the mIoU decreases between 2.4 % (Xception-41) and 7.1 % (MobileNet-V2). Therefore, the corresponding CD scores are oftentimes considerably high. On PASCA...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.