REVIEW 22 cited by
AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty
read the original abstract
Modern deep neural networks can achieve high accuracy when the training distribution and test distribution are identically distributed, but this assumption is frequently violated in practice. When the train and test distributions are mismatched, accuracy can plummet. Currently there are few techniques that improve robustness to unforeseen data shifts encountered during deployment. In this work, we propose a technique to improve the robustness and uncertainty estimates of image classifiers. We propose AugMix, a data processing technique that is simple to implement, adds limited computational overhead, and helps models withstand unforeseen corruptions. AugMix significantly improves robustness and uncertainty measures on challenging image classification benchmarks, closing the gap between previous methods and the best possible performance in some cases by more than half.
Forward citations
Cited by 22 Pith papers
-
Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization
PAPT uses adversarial prompt tuning on diffusion models to generate domain-style images while preserving category features, claiming superior single-domain generalization performance.
-
Adaptive Camera Sensor for Vision Models
Lens adapts camera sensors in real time via the VisiT confidence-based quality indicator to improve vision model accuracy on domain-shifted images, shown on ImageNet-ES and a new diverse benchmark.
-
Radial Basis Function Networks as Projection Heads in Self-Supervised Learning
RBFN projection heads serve as competitive replacements for MLP heads in SSL and enable SNS, a label-free metric from RBF parameters that correlates strongly with logistic regression evaluation.
-
Learning from almost nothing: How neural networks survive heavy input corruption
Infinite-width MLPs implement a nearest-class-mean prototype classifier as their leading-order decision rule under heavy attribute noise, explaining observed robustness in experiments.
-
ReSAGE-PAR: Representational Similarity Assessment for Generative Expansion in Pedestrian Attribute Recognition
A LoRA-adapted diffusion model generates surveillance-style pedestrian images whose labels are verified by a Bayesian classifier on BLIP scores, improving PAR accuracy over naive synthetic labeling.
-
On the Reliability of Cue Conflict and Beyond
Stylized cue-conflict bias scores are confounded by impure cues, imbalance, ratio metrics and restricted labels; REFINED-BIAS supplies pure balanced stimuli and full-label MRR sensitivity for reliable diagnosis.
-
ReFace: Reorganizing Facial Spatiotemporal Representations for Improved Pain Assessment
Reorganizing facial video into four channel-concatenated quadrants before tokenization yields 56.00% test accuracy on AI4Pain video-only pain classification, the highest reported under that benchmark protocol.
-
Interleaved Noise Injection Improves Clean, Corrupted, and OOD Performance
Interleaving clean and noisy training epochs improves clean, corrupted, and out-of-distribution accuracy on CIFAR-100 and ImageNet for CNNs and ViTs, with impulse noise best for ResNets and Gaussian noise best for ViTs.
-
ReSAGE-PAR: Representational Similarity Assessment for Generative Expansion in Pedestrian Attribute Recognition
ReSAGE-PAR adapts diffusion models with LoRA, scores generated images via vision-language prompts, and applies Bayesian classification to produce pseudo-labels, yielding up to 8.7% gains when used to expand PAR datasets.
-
Dual Feature Decoupling for Fine-Grained OOD Detection
DFDNet disentangles content from style via dual modules to boost fine-grained OOD detection performance on multiple datasets.
-
FDDet: Achieving Data-Efficient Food Defect Detection Under Real-World Scenarios
FDDet is a semi-supervised object detection framework with BBoxMixUp and CGPC that outperforms standard detectors on the new FDD-48 food defect dataset under data-limited real-world conditions.
-
A Composite Activation Function for Learning Stable Binary Representations
HTAF is a sigmoid-tanh composite that approximates the Heaviside function to allow stable gradient training of binary activation networks, yielding ICBMs with stable discretization and competitive performance on image tasks.
-
TINS: Test-time ID-prototype-separated Negative Semantics Learning for OOD Detection
TINS improves OOD detection by learning negative semantics at test time with ID-prototype separation, cutting average FPR95 from 14.04% to 6.72% on the Four-OOD benchmark with ImageNet-1K.
-
Medical Model Synthesis Architectures: A Case Study
MedMSA framework retrieves knowledge via language models then builds formal probabilistic models to produce uncertainty-weighted differential diagnoses from symptoms.
-
Agentic AIs Are the Missing Paradigm for Out-of-Distribution Generalization in Foundation Models
Agentic AI systems are required to overcome the parameter coverage ceiling that prevents foundation models from handling certain out-of-distribution cases.
-
WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval
WRF4CIR uses weight-regularized fine-tuning with adversarial perturbations to mitigate overfitting in composed image retrieval and narrows the generalization gap on benchmarks.
-
A Patch-based Cross-view Regularized Framework for Backdoor Defense in Multimodal Large Language Models
A patch-augmented cross-view regularization method reduces backdoor attack success rates in multimodal LLMs by enforcing output differences between original and perturbed views while using entropy constraints to prese...
-
Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method
A progressive curriculum that trains open-vocabulary detectors on low-ambiguity, high-signal cross-modal alignments first improves robustness to visual domain shifts, with modest, test-tuned gains.
-
Data-Augmented Quantization-Aware Knowledge Distillation
A teacher-only metric, M=DEV-CMI, ranks data augmentations for low-bit quantized knowledge distillation and improves accuracy on CIFAR and Tiny ImageNet.
-
A Unified Tokenization Framework for Pain Recognition using Heterogeneous 3D Modalities
A unified tokenizer maps facial video and fNIRS into one token space; the segment-latent transformer hits 57.33% test accuracy on AI4Pain pain recognition.
-
Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution
DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...
-
Improving Facial Emotion Recognition through Dataset Merging and Balanced Training Strategies
Merging CK+, FER+, and KDEF datasets with online/offline augmentation and random weighted sampling enables a deep CNN to classify seven facial emotions at 82% accuracy.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.