Pith. sign in

REVIEW 1 cited by

ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic Object

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.18775 v1 pith:3YDTHW7G submitted 2024-03-27 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords modelsrobustnesssyntheticdiffusionimagenet-dimagesworkaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We establish rigorous benchmarks for visual perception robustness. Synthetic images such as ImageNet-C, ImageNet-9, and Stylized ImageNet provide specific type of evaluation over synthetic corruptions, backgrounds, and textures, yet those robustness benchmarks are restricted in specified variations and have low synthetic quality. In this work, we introduce generative model as a data source for synthesizing hard images that benchmark deep models' robustness. Leveraging diffusion models, we are able to generate images with more diversified backgrounds, textures, and materials than any prior work, where we term this benchmark as ImageNet-D. Experimental results show that ImageNet-D results in a significant accuracy drop to a range of vision models, from the standard ResNet visual classifier to the latest foundation models like CLIP and MiniGPT-4, significantly reducing their accuracy by up to 60\%. Our work suggests that diffusion models can be an effective source to test vision models. The code and dataset are available at https://github.com/chenshuang-zhang/imagenet_d.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vision Transformer Neural Architecture Search for Out-of-Distribution Generalization: Benchmark and Insights

    cs.LG 2025-01 conditional novelty 7.0 of 10

    A 3000-architecture benchmark shows ViT OoD accuracy varies widely with architecture and that embedding dimension is the strongest structural correlate of OoD robustness.

Pith tools