Pith. sign in

REVIEW 3 cited by

Dataset Quantization with Active Learning based Adaptive Sampling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.07268 v1 pith:PF7P2P7U submitted 2024-07-09 cs.CV

classification cs.CV
keywords datasetquantizationclasseslearningperformancesamplingactiveadaptive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep learning has made remarkable progress recently, largely due to the availability of large, well-labeled datasets. However, the training on such datasets elevates costs and computational demands. To address this, various techniques like coreset selection, dataset distillation, and dataset quantization have been explored in the literature. Unlike traditional techniques that depend on uniform sample distributions across different classes, our research demonstrates that maintaining performance is feasible even with uneven distributions. We find that for certain classes, the variation in sample quantity has a minimal impact on performance. Inspired by this observation, an intuitive idea is to reduce the number of samples for stable classes and increase the number of samples for sensitive classes to achieve a better performance with the same sampling ratio. Then the question arises: how can we adaptively select samples from a dataset to achieve optimal performance? In this paper, we propose a novel active learning based adaptive sampling strategy, Dataset Quantization with Active Learning based Adaptive Sampling (DQAS), to optimize the sample selection. In addition, we introduce a novel pipeline for dataset quantization, utilizing feature space from the final stage of dataset quantization to generate more precise dataset bins. Our comprehensive evaluations on the multiple datasets show that our approach outperforms the state-of-the-art dataset compression methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A computational fluid dynamics model for the simulation of flashboiling flow inside pressurized metered dose inhalers

    physics.flu-dyn 2025-08 unverdicted novelty 6.0 of 10

    The abstract claims a first-of-kind open-source CFD model, combining Volume-of-Fluid and cavitation modeling, that quantitatively predicts flashboiling flow in pressurized metered dose inhalers, but the supplied text ...

  2. CaO$_2$: Rectifying Inconsistencies in Diffusion-Based Dataset Distillation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CaO2 selects confident diffusion-generated samples and optimizes their latents against the denoising objective, achieving state-of-the-art distilled-dataset accuracy on ImageNet subsets.

  3. NMS: Efficient Edge DNN Training via Near-Memory Sampling on Manifolds

    cs.LG 2025-08 reject novelty 5.0 of 10

    A t-SNE-based sampling algorithm with differential evolution and near-memory hardware is claimed to speed up edge DNN training and reduce memory energy.

Pith tools