Pith. sign in

REVIEW 3 cited by

Context-guided Responsible Data Augmentation with Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.10687 v1 pith:JECON62U submitted 2025-03-12 cs.CV

classification cs.CV
keywords generativeaugmentationdatadiffusionimagesmodelmodelsnatural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Generative diffusion models offer a natural choice for data augmentation when training complex vision models. However, ensuring reliability of their generative content as augmentation samples remains an open challenge. Despite a number of techniques utilizing generative images to strengthen model training, it remains unclear how to utilize the combination of natural and generative images as a rich supervisory signal for effective model induction. In this regard, we propose a text-to-image (T2I) data augmentation method, named DiffCoRe-Mix, that computes a set of generative counterparts for a training sample with an explicitly constrained diffusion model that leverages sample-based context and negative prompting for a reliable augmentation sample generation. To preserve key semantic axes, we also filter out undesired generative samples in our augmentation process. To that end, we propose a hard-cosine filtration in the embedding space of CLIP. Our approach systematically mixes the natural and generative images at pixel and patch levels. We extensively evaluate our technique on ImageNet-1K,Tiny ImageNet-200, CIFAR-100, Flowers102, CUB-Birds, Stanford Cars, and Caltech datasets, demonstrating a notable increase in performance across the board, achieving up to $\sim 3\%$ absolute gain for top-1 accuracy over the state-of-the-art methods, while showing comparable computational overhead. Our code is publicly available at https://github.com/khawar-islam/DiffCoRe-Mix

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition

    cs.CV 2026-07 conditional novelty 6.0 of 10

    On vein-recognition benchmarks, mixup-style augmentations win on clean accuracy but hurt calibration and adversarial robustness, while simple geometric transforms usually hurt performance.

  2. SimDiffRec: Semantic Similarity-Guided Diffusion for Contrastive Sequential Recommendation

    cs.IR 2025-07 conditional novelty 5.0 of 10

    SimDiffRec augments user sequences by replacing items at high-confidence diffusion positions with the model's top prediction and using averaged similar-item embeddings as noise, reporting consistent but unverified gai...

  3. Emerging AI Approaches for Cancer Spatial Omics

    q-bio.QM 2025-06 unverdicted novelty 2.0 of 10

    A review that groups AI methods for cancer spatial omics into data-driven, constraint-based, and mechanistic modeling paradigms, calling for more interpretable models and mouse-model-generated perturbational data.

Pith tools