REVIEW 5 cited by
Feedback-guided Data Synthesis for Imbalanced Classification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Current status quo in machine learning is to use static datasets of real images for training, which often come from long-tailed distributions. With the recent advances in generative models, researchers have started augmenting these static datasets with synthetic data, reporting moderate performance improvements on classification tasks. We hypothesize that these performance gains are limited by the lack of feedback from the classifier to the generative model, which would promote the usefulness of the generated samples to improve the classifier's performance. In this work, we introduce a framework for augmenting static datasets with useful synthetic samples, which leverages one-shot feedback from the classifier to drive the sampling of the generative model. In order for the framework to be effective, we find that the samples must be close to the support of the real data of the task at hand, and be sufficiently diverse. We validate three feedback criteria on a long-tailed dataset (ImageNet-LT) as well as a group-imbalanced dataset (NICO++). On ImageNet-LT, we achieve state-of-the-art results, with over 4 percent improvement on underrepresented classes while being twice efficient in terms of the number of generated synthetic samples. NICO++ also enjoys marked boosts of over 5 percent in worst group accuracy. With these results, our framework paves the path towards effectively leveraging state-of-the-art text-to-image models as data sources that can be queried to improve downstream applications.
Forward citations
Cited by 5 Pith papers
-
Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting
Post-generation selection via Homogeneous-Heterogeneous real-data splits and a fidelity-diversity score raises synthetic-image utility for classification and segmentation without retraining generators.
-
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
Per-image LoRA adapters fused at inference time produce synthetic training data that improves few-shot image classification accuracy over existing synthetic-data methods.
-
Boosting Statistic Learning with Synthetic Data from Pretrained Large Models
The paper claims synthetic tabular data generated by pass-through Stable Diffusion, filtered by Wasserstein distance or hypothesis tests, improves predictive accuracy, but the evidence is weakened by missing baselines...
-
Gaussian Splatting is an Effective Data Generator for 3D Object Detection
Inserting Gaussian-splat reconstructed 3D objects into reconstructed driving scenes is a more effective augmentation for camera-based 3D object detection than diffusion-based image synthesis.
-
The Pitfalls of Memorization: When Memorization Hurts Generalization
Memorization-aware training (MAT) shifts logits using calibrated held-out predictions from an XRM auxiliary model, improving worst-group accuracy under subpopulation shift.
Discussion (0). Continue with ORCID to comment.