REVIEW 15 cited by
Image Data Augmentation for Deep Learning: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Deep learning has achieved remarkable results in many computer vision tasks. Deep neural networks typically rely on large amounts of training data to avoid overfitting. However, labeled data for real-world applications may be limited. By improving the quantity and diversity of training data, data augmentation has become an inevitable part of deep learning model training with image data. As an effective way to improve the sufficiency and diversity of training data, data augmentation has become a necessary part of successful application of deep learning models on image data. In this paper, we systematically review different image data augmentation methods. We propose a taxonomy of reviewed methods and present the strengths and limitations of these methods. We also conduct extensive experiments with various data augmentation methods on three typical computer vision tasks, including semantic segmentation, image classification and object detection. Finally, we discuss current challenges faced by data augmentation and future research directions to put forward some useful research guidance.
Forward citations
Cited by 15 Pith papers
-
Improving Backward Conformal Prediction via Non-Conformity Score Transformation
ST-BCP tightens the coverage bound in Backward Conformal Prediction by applying a computable data-dependent transformation to nonconformity scores, reducing the average gap from 4.20% to 1.12% on benchmarks while prov...
-
Beyond and Free from Diffusion: Invertible Guided Consistency Training
iGCT trains guided consistency models from scratch by mixing the original noise with a direction to a random target-class image, and reports better FID and precision than classifier-free guidance at high guidance on CIFAR-10.
-
Lights, Camera, Malfunction: When Illumination Robustness Leaves VLA Models Blind to Color
Aggressive color data augmentation makes VLA robot models lighting-robust by teaching them to ignore color, causing failure on benign color-dependent tasks; fixing hue during adversarial training (ChromaGuard) preserves both.
-
Learning from Noise: Enhancing DNNs for Event-Based Vision through Controlled Noise Injection
Adding synthetic shot noise at random intensities to event-camera training data makes CNN, ViT, SNN, and GCN classifiers robust to input noise, outperforming test-time filtering.
-
Image Quality Dependent Degradation for AI Systems
A normalizing-flow quality monitor that lowers an object detector's confidence threshold on low-quality images raises pedestrian recall by a few points while slightly reducing precision.
-
Assessing the Operational Impact of Poisoning Attacks over Augmented 3D Point Cloud Public Datasets for Connected and Autonomous Vehicles
GAN-based augmentation of poisoned 3D point cloud datasets amplifies attack effectiveness, increasing misclassification and operational impact on CAV decision-making by up to 3x compared to non-augmented baselines.
-
NoiseCutMix: A Novel Data Augmentation Approach by Mixing Estimated Noise in Diffusion Models
Mixing the estimated noise of two class prompts at each denoising step of Stable Diffusion generates natural augmented images that improve fine-grained classification over CutMix on some datasets.
-
Group Relative Augmentation for Data Efficient Action Detection
A LoRA plus FiLM feature-augmentation method with a group-weighted loss reports modest few-shot action detection gains on AVA and MOMA, but the evidence for the weighting component is weak.
-
Curvature Enhanced Data Augmentation for Regression
CEMS augments regression training by sampling from a second-order, curvature-aware local model of the joint input-output manifold, and reports competitive in-distribution and out-of-distribution results on nine benchmarks.
-
Automatic detection of overshooting tops and their properties from visible satellite channels
A CNN trained on about 10,000 manually labeled European cases detects overshooting tops from visible satellite images with 97.7% probability of detection and estimates their height with a mean error of about 0.25 km.
-
AI for Cultural Heritage Textiles: Fine-Tuned Latent Diffusion for Novel Ulos Motif Synthesis
Fine-tuned Protogen v3.4 generates novel Ulos motifs with ~10.5 imes lower FID and 2× higher IS than Stable Diffusion v1.4, with guidance scale 5–9 balancing fidelity and diversity.
-
Hybrid Ensemble Approaches: Optimal Deep Feature Fusion and Hyperparameter-Tuned Classifier Ensembling for Enhanced Brain Tumor Classification
A double ensemble that fuses features from pretrained CNNs and ViTs and ensembles tuned ML classifiers reaches 97.5% to 99.3% accuracy on three public brain MRI datasets, but the gains are not benchmarked against a he...
-
Hierarchical Deep Feature Fusion and Ensemble Learning for Enhanced Brain Tumor MRI Classification
A ViT feature ensemble plus ML classifier voting pipeline is evaluated on two binary brain MRI datasets, reporting up to 99.8% accuracy without a same-dataset comparison against prior methods.
-
Systematic Integration of Attention Modules into CNNs for Accurate and Generalizable Medical Image Diagnosis
Attention-augmented CNNs usually beat plain CNNs on two medical image datasets, with EfficientNetB5 plus hybrid attention the best, but test-set-based model selection undermines the claimed consistency.
-
3D Skeleton-Based Action Recognition: A Review
A task-oriented review of skeleton-based action recognition that reorganizes known methods along a data processing pipeline and contains no new experimental result.
Discussion (0). Continue with ORCID to comment.