Pith. sign in

REVIEW 6 cited by

Natural Adversarial Examples

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.07174 v4 pith:PKVCDR33 submitted 2019-07-16 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords datasetsmodelsadversarialdatasetdetectionout-of-distributionperformanceaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce two challenging datasets that reliably cause machine learning model performance to substantially degrade. The datasets are collected with a simple adversarial filtration technique to create datasets with limited spurious cues. Our datasets' real-world, unmodified examples transfer to various unseen models reliably, demonstrating that computer vision models have shared weaknesses. The first dataset is called ImageNet-A and is like the ImageNet test set, but it is far more challenging for existing models. We also curate an adversarial out-of-distribution detection dataset called ImageNet-O, which is the first out-of-distribution detection dataset created for ImageNet models. On ImageNet-A a DenseNet-121 obtains around 2% accuracy, an accuracy drop of approximately 90%, and its out-of-distribution detection performance on ImageNet-O is near random chance levels. We find that existing data augmentation techniques hardly boost performance, and using other public training datasets provides improvements that are limited. However, we find that improvements to computer vision architectures provide a promising path towards robust models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Certified Circuits: Stability Guarantees for Mechanistic Circuits

    cs.AI 2026-02 conditional novelty 6.0 of 10

    Certified Circuits uses deletion-based randomized smoothing to guarantee that circuit components stay included or excluded under bounded edits to the concept dataset, yielding more compact and more accurate circuits.

  2. Improving Detection of Rare Nodes in Hierarchical Multi-Label Learning

    cs.LG 2026-02 conditional novelty 6.0 of 10

    A node-weighted loss combining inverse-frequency weighting and ensemble-uncertainty focal terms improves recall of rare classes in hierarchical multi-label models by up to ~5x.

  3. Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Dense scaling-law fits across model sizes, datasets and tasks show that MaMMUT (contrastive plus captioning loss) outperforms standard CLIP at large compute scales, with a consistent crossover around 1e10 to 1e11 GFLOPs.

  4. Domain Adaptation via Feature Refinement

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    DAFR2 combines target-data batch normalization adaptation, feature distillation, and hypothesis transfer to make models robust to image corruption without target labels.

  5. Quality over Quantity: An Effective Large-Scale Data Reduction Strategy Based on Pointwise V-Information

    cs.LG 2025-06 reject novelty 4.0 of 10

    A PVI-based data reduction and progressive training strategy is applied to Chinese NLI, but the reported small accuracy declines do not match the experimental tables.

  6. Revisiting Bayesian Model Averaging in the Era of Foundation Models

    cs.LG 2025-05 reject novelty 4.0 of 10

    The paper proposes Bayesian model averaging and an entropy-minimizing weight optimizer for ensembling foundation models, reporting accuracy gains over output averaging on image and text classification tasks.

Pith tools