REVIEW 3 cited by
Why do deep convolutional networks generalize so poorly to small image transformations?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Convolutional Neural Networks (CNNs) are commonly assumed to be invariant to small image transformations: either because of the convolutional architecture or because they were trained using data augmentation. Recently, several authors have shown that this is not the case: small translations or rescalings of the input image can drastically change the network's prediction. In this paper, we quantify this phenomena and ask why neither the convolutional architecture nor data augmentation are sufficient to achieve the desired invariance. Specifically, we show that the convolutional architecture does not give invariance since architectures ignore the classical sampling theorem, and data augmentation does not give invariance because the CNNs learn to be invariant to transformations only for images that are very similar to typical images from the training set. We discuss two possible solutions to this problem: (1) antialiasing the intermediate representations and (2) increasing data augmentation and show that they provide only a partial solution at best. Taken together, our results indicate that the problem of insuring invariance to small image transformations in neural networks while preserving high accuracy remains unsolved.
Forward citations
Cited by 3 Pith papers
-
Benchmarking the Robustness of Semantic Segmentation Models
A large-scale benchmark of semantic segmentation models under 19 image corruptions shows that, within DeepLabv3+, better clean-data performance usually comes with better corruption robustness, while Dense Prediction C...
-
SNAP: A Benchmark for Testing the Effects of Capture Conditions on Fundamental Vision Tasks
A new benchmark shows that camera capture settings and lighting systematically change the performance of image classifiers, object detectors, and VQA models, and that common vision datasets are biased toward narrow ex...
-
A Possible Reason for why Data-Driven Beats Theory-Driven Computer Vision
Deep learning may have beaten classical computer vision partly because benchmarks used camera settings outside the classical algorithms' operating ranges.
Discussion (0). Continue with ORCID to comment.