Pith. sign in

REVIEW 22 cited by

AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.02781 v2 pith:LLQHMVSO submitted 2019-12-05 stat.ML cs.CVcs.LG

AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty

classification stat.ML cs.CVcs.LG
keywords robustnessaugmixdataimproveuncertaintyaccuracydistributionimage
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Modern deep neural networks can achieve high accuracy when the training distribution and test distribution are identically distributed, but this assumption is frequently violated in practice. When the train and test distributions are mismatched, accuracy can plummet. Currently there are few techniques that improve robustness to unforeseen data shifts encountered during deployment. In this work, we propose a technique to improve the robustness and uncertainty estimates of image classifiers. We propose AugMix, a data processing technique that is simple to implement, adds limited computational overhead, and helps models withstand unforeseen corruptions. AugMix significantly improves robustness and uncertainty measures on challenging image classification benchmarks, closing the gap between previous methods and the best possible performance in some cases by more than half.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization

    cs.CV 2026-06 unverdicted novelty 7.0

    PAPT uses adversarial prompt tuning on diffusion models to generate domain-style images while preserving category features, claiming superior single-domain generalization performance.

  2. Adaptive Camera Sensor for Vision Models

    cs.CV 2025-03 unverdicted novelty 7.0

    Lens adapts camera sensors in real time via the VisiT confidence-based quality indicator to improve vision model accuracy on domain-shifted images, shown on ImageNet-ES and a new diverse benchmark.

  3. Radial Basis Function Networks as Projection Heads in Self-Supervised Learning

    cs.CV 2026-06 unverdicted novelty 6.0

    RBFN projection heads serve as competitive replacements for MLP heads in SSL and enable SNS, a label-free metric from RBF parameters that correlates strongly with logistic regression evaluation.

  4. Learning from almost nothing: How neural networks survive heavy input corruption

    cs.LG 2026-06 unverdicted novelty 6.0

    Infinite-width MLPs implement a nearest-class-mean prototype classifier as their leading-order decision rule under heavy attribute noise, explaining observed robustness in experiments.

  5. ReSAGE-PAR: Representational Similarity Assessment for Generative Expansion in Pedestrian Attribute Recognition

    cs.CV 2026-06 conditional novelty 6.0

    A LoRA-adapted diffusion model generates surveillance-style pedestrian images whose labels are verified by a Bayesian classifier on BLIP scores, improving PAR accuracy over naive synthetic labeling.

  6. On the Reliability of Cue Conflict and Beyond

    cs.CV 2026-03 conditional novelty 6.0

    Stylized cue-conflict bias scores are confounded by impure cues, imbalance, ratio metrics and restricted labels; REFINED-BIAS supplies pure balanced stimuli and full-label MRR sensitivity for reliable diagnosis.

  7. ReFace: Reorganizing Facial Spatiotemporal Representations for Improved Pain Assessment

    cs.CV 2026-07 conditional novelty 5.0

    Reorganizing facial video into four channel-concatenated quadrants before tokenization yields 56.00% test accuracy on AI4Pain video-only pain classification, the highest reported under that benchmark protocol.

  8. Interleaved Noise Injection Improves Clean, Corrupted, and OOD Performance

    cs.LG 2026-07 conditional novelty 5.0

    Interleaving clean and noisy training epochs improves clean, corrupted, and out-of-distribution accuracy on CIFAR-100 and ImageNet for CNNs and ViTs, with impulse noise best for ResNets and Gaussian noise best for ViTs.

  9. ReSAGE-PAR: Representational Similarity Assessment for Generative Expansion in Pedestrian Attribute Recognition

    cs.CV 2026-06 unverdicted novelty 5.0

    ReSAGE-PAR adapts diffusion models with LoRA, scores generated images via vision-language prompts, and applies Bayesian classification to produce pseudo-labels, yielding up to 8.7% gains when used to expand PAR datasets.

  10. Dual Feature Decoupling for Fine-Grained OOD Detection

    cs.CV 2026-06 unverdicted novelty 5.0

    DFDNet disentangles content from style via dual modules to boost fine-grained OOD detection performance on multiple datasets.

  11. FDDet: Achieving Data-Efficient Food Defect Detection Under Real-World Scenarios

    cs.CV 2026-05 unverdicted novelty 5.0

    FDDet is a semi-supervised object detection framework with BBoxMixUp and CGPC that outperforms standard detectors on the new FDD-48 food defect dataset under data-limited real-world conditions.

  12. A Composite Activation Function for Learning Stable Binary Representations

    cs.LG 2026-05 unverdicted novelty 5.0

    HTAF is a sigmoid-tanh composite that approximates the Heaviside function to allow stable gradient training of binary activation networks, yielding ICBMs with stable discretization and competitive performance on image tasks.

  13. TINS: Test-time ID-prototype-separated Negative Semantics Learning for OOD Detection

    cs.CV 2026-05 unverdicted novelty 5.0

    TINS improves OOD detection by learning negative semantics at test time with ID-prototype separation, cutting average FPR95 from 14.04% to 6.72% on the Four-OOD benchmark with ImageNet-1K.

  14. Medical Model Synthesis Architectures: A Case Study

    cs.AI 2026-05 unverdicted novelty 5.0

    MedMSA framework retrieves knowledge via language models then builds formal probabilistic models to produce uncertainty-weighted differential diagnoses from symptoms.

  15. Agentic AIs Are the Missing Paradigm for Out-of-Distribution Generalization in Foundation Models

    cs.LG 2026-05 unverdicted novelty 5.0

    Agentic AI systems are required to overcome the parameter coverage ceiling that prevents foundation models from handling certain out-of-distribution cases.

  16. WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval

    cs.CV 2026-04 unverdicted novelty 5.0

    WRF4CIR uses weight-regularized fine-tuning with adversarial perturbations to mitigate overfitting in composed image retrieval and narrows the generalization gap on benchmarks.

  17. A Patch-based Cross-view Regularized Framework for Backdoor Defense in Multimodal Large Language Models

    cs.CV 2026-04 unverdicted novelty 5.0

    A patch-augmented cross-view regularization method reduces backdoor attack success rates in multimodal LLMs by enforcing output differences between original and perturbed views while using entropy constraints to prese...

  18. Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method

    cs.CV 2026-03 conditional novelty 5.0

    A progressive curriculum that trains open-vocabulary detectors on low-ambiguity, high-signal cross-modal alignments first improves robustness to visual domain shifts, with modest, test-tuned gains.

  19. Data-Augmented Quantization-Aware Knowledge Distillation

    cs.LG 2025-09 conditional novelty 5.0

    A teacher-only metric, M=DEV-CMI, ranks data augmentations for low-bit quantized knowledge distillation and improves accuracy on CIFAR and Tiny ImageNet.

  20. A Unified Tokenization Framework for Pain Recognition using Heterogeneous 3D Modalities

    cs.CV 2026-07 conditional novelty 4.0

    A unified tokenizer maps facial video and fNIRS into one token space; the segment-latent transformer hits 57.33% test accuracy on AI4Pain pain recognition.

  21. Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution

    cs.CV 2026-07 reject novelty 4.0

    DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...

  22. Improving Facial Emotion Recognition through Dataset Merging and Balanced Training Strategies

    cs.CV 2026-04 unverdicted novelty 2.0

    Merging CK+, FER+, and KDEF datasets with online/offline augmentation and random weighted sampling enables a deep CNN to classify seven facial emotions at 82% accuracy.