REVIEW 5 major objections 6 minor 40 references
Adaptive Dataset Quantization
T0 review · 5 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Scoring data bins by texture and diversity beats uniform sampling
desk verdict The adaptive sampling idea is plausible, but Eq. 9–10 don't enforce a fixed keep ratio, so the reported gains over DQ are uncontrolled until the sampling step is clarified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the per-bin importance score $\hat{I}_n = \mathrm{Norm}(\mathrm{Rep}_n) + \mathrm{Norm}(\mathrm{Div}_n)$, where $\mathrm{Rep}_n$ is the mean gradient magnitude of a bin's $L\times L$ patches, $T(P) = \frac{1}{L^2}\sum_{i,j} G(P_{i,j})$, and $\mathrm{Div}_n$ is the negative contrastive ratio from a discriminator. The score feeds the adaptive sampling ratio $r_n = \alpha \hat{I}_n + (1-\alpha) N(n)/\sum_n N(n)$, replacing uniform sampling. The paper's argument stands on this scalar correctly ranking bins by their contribution to training; the texture-level identity $\mathrm{Rep} = T$ is justified by a single trajectory-distance experiment and is never derived from first principles.
What would settle it
Train a model on a compressed set made with ADQ's importance scores, and on a second set where the same scores are randomly permuted across bins before sampling; if accuracy does not drop, the scores are not the active ingredient. A direct check is to compute, on ImageNet-1K with the ViT feature extractor, the rank correlation between each bin's texture level and the actual parameter-space distance between a student trained on that bin and an expert trained on the full dataset; a non-monotonic or near-zero correlation would falsify the $\mathrm{Rep}=T$ mapping.
Extended reading notes
Core claim
The central claim is that the bins produced by DQ's recursive submodular selection differ systematically in how much they contribute to training, so the uniform sampling used by naive DQ is a spurious balance rather than an optimal one. ADQ makes this difference measurable by assigning each bin a representativeness score $\mathrm{Rep}_n$ equal to its texture level $T(P)$, a diversity score $\mathrm{Div}_n$ from a contrastive instance-discrimination discriminator, and an importance score $\hat{I}_n = \mathrm{Norm}(\mathrm{Rep}_n) + \mathrm{Norm}(\mathrm{Div}_n)$. Sampling then uses the proportion $r_n = \alpha \hat{I}_n + (1-\alpha)\, N(n)/\sum_{n} N(n)$ for each bin. The authors show that this adaptive sampling outperforms uniform sampling on every architecture they test, with the largest relative gains at higher keep ratios, and that the extra computation is near negligible.
Load-bearing premise
The load-bearing premise is that the texture level of a bin's patches measures how representative that bin is of the full dataset; this is supported by one trajectory-distance curve on ResNet-18/CIFAR-10, with no verification on the other architectures, datasets, or bin orders used in the main experiments.
Editorial extensions
If this is right
- Any DQ-based pipeline can adopt ADQ as a drop-in replacement for uniform sampling, gaining 2–4 percentage points on CIFAR-10 without changing bin generation or downstream training.
- At a 60% data keep ratio, the compressed set matches full-data accuracy, so storage and bandwidth can be cut by 40% with no accuracy loss.
- The improvement transfers across CNNs and transformers, so the method is not tied to a single architecture family.
- The full generation cost stays close to DQ's, making the accuracy gain available at roughly the same compute budget.
Reading between the lines
- Not tested in the paper: the same adaptive sampling could be layered on top of any bin-generation scheme, not only DQ's GraphCut bins, since the scores are computed after bins exist.
- The equal-weight sum of normalized RS and DS is a design choice; a multiplicative combination or a learned weight might do better, especially at high keep ratios where the paper observes the optimal $\alpha$ shifting.
- The texture-level representativeness proxy is a natural-image heuristic; extending ADQ to non-image data would require a different cheap representativeness measure, which the paper does not propose.
- A harder stress test, not in the paper, would be full-scale ImageNet-21k, where the single trajectory-distance justification for texture level is under less direct control.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Adaptive Dataset Quantization (ADQ), an extension of Dataset Quantization (DQ) that replaces uniform sampling across bins with an adaptive sampling strategy. Each bin is scored by a representativeness score based on texture level, a diversity score based on contrastive learning, and an importance score defined as the sum of the normalized RS and DS. Adaptive sampling then draws more images from bins with higher importance. The authors report consistent accuracy gains over DQ on CIFAR-10 across five architectures (2.6%, 2.8%, and 3.3% average improvements at 10%, 20%, and 30% keep ratios), similar improvements on ImageNet-1K, and lossless compression at 60% data, with negligible additional generation time.
Significance. If the reported results hold, ADQ would be a practical, low-cost improvement over DQ that preserves cross-architecture generalization. The paper ships code, reports five-run averages with error bars, and identifies a genuine limitation of DQ's uniform sampling. However, the central quantitative claim is not yet established because the method as written does not control the size of the compressed set, and several components of the scoring mechanism are underspecified. The strengths are real but contingent on fixing the sampling-size issue.
major comments (5)
- [Adaptive Sampling, Eqs. (9)–(10); Table 1] The method does not enforce the target keep ratio ρ. With m equal bins of size K=M/m, Eq. (9) gives r_n = α Î_n + (1−α)/m and Eq. (10) gives q_n = floor(r_n K), so the total selected size is approximately (M/m)[α Σ_n Î_n + (1−α)]. If Norm in Eqs. (6)–(7) is z-score normalization, ΣÎ_n = 0 and the total is about (1−α)M/m, which for m=10 and α=0.6 is 4% of the dataset rather than the nominal 10%. If Norm is min-max normalization, the total varies with the data and can range far above or below ρM. The floor function only bounds each bin's contribution and cannot make Σ q_n equal the intended ρM. Therefore the reported gains over DQ may be an artifact of training on a different number of images rather than of adaptive sampling. The authors must specify the normalization, add an explicit rescaling or completion step to match the target size, and rerun the comparisons at matched dataset sizes.
- [Representativeness Score, Eq. (3) and Fig. 4] The identification Rep(·)=T(·) is supported only by a single experiment comparing three global texture-level batches on ResNet-18/CIFAR-10. This does not establish that mean gradient magnitude ranks individual DQ bins by their trajectory closeness to the expert model, nor that the relationship transfers to other architectures, datasets, or bin orders. Since this ordering directly determines sampling proportions, the mechanism remains a heuristic. A per-dataset validation at the bin level, or a derivation, is needed before the representativeness score can be considered a principled component of the method.
- [Diversity Score, Eq. (5)] Eq. (5) is mathematically incomplete. The term N(x−) is not defined as a function of x_i even though it appears inside an expectation over x_i, and the leading negative sign is unexplained. As written, Div(S_n) is not a standard contrastive objective, and it is unclear whether larger or smaller values correspond to greater diversity. This ambiguity propagates into the importance score in Eq. (8) and into the adaptive sampling in Eq. (9).
- [Ablation Study, Fig. 6 and Comparisons with Previous Methods] The paper states that 'During the actual experiment, we adjust corresponding values of α according to different architectures,' but it does not report the chosen α values or the selection criterion. If α is selected based on validation accuracy for each architecture and each keep ratio, while the DQ baseline has no analogous tuned parameter, the comparison in Table 1 and Fig. 5 is not apples-to-apples. The authors should report the α values used for every architecture and dataset, and ideally show results with a fixed α in addition to tuned α.
- [Comparisons with Previous Methods] The paper claims that 'our ADQ also achieves lossless results with only 60% of the data' on CIFAR-10 and ImageNet-1K, but no 60% results are tabulated and no criterion for 'lossless' is given. Please provide the numerical results with error bars at ρ=60% and specify the statistical threshold used to declare lossless compression.
minor comments (6)
- [Importance Evaluation, Eqs. (6)–(7)] The normalization operator Norm is never specified (z-score, min-max, or other). The citation to Ioffe and Szegedy (2015) is for batch normalization, which is not a generic normalization method; please define the operation explicitly.
- [Related Works, Remark] There is a typo: 'excessively comrplex' should be 'excessively complex'.
- [Experimental Setup, Datasets] The ImageNet-1K training size is given as '128,1126 samples', which appears to be a typo; the standard value is 1,281,167.
- [Figure 2] The subplot axes are unlabeled and the legend text is partially garbled, which makes the claimed trends in RS, DS, and IS difficult to read. Please provide clear axis labels and legends.
- [Algorithm 1] The input list 'Required:' is unconventional; the hyperparameters L, τ, and α should be listed as algorithm parameters rather than required inputs, and the output should note that the patch-dropping stage is applied after the initial compressed dataset is formed.
- [Adaptive Sampling, Eq. (9)] The equal weighting of RS and DS in Eq. (8) is asserted rather than derived. The ablation study reports that DS contributes slightly more than RS, so a short discussion of why equal weights are preferred across datasets would improve the presentation.
Circularity Check
No significant circularity: ADQ's importance score is a hand-defined heuristic, the accuracy claims are empirical comparisons against external baselines, and there are no load-bearing self-citations.
full rationale
ADQ is an empirical method layered on DQ. The representativeness score is stipulated as texture level (Rep(·)=T(·), Eq. 3) and justified by the trajectory-distance experiment in Fig. 4; the diversity score is borrowed from contrastive learning; and the importance score is defined by Eq. 8 as the normalized sum of the two. Using that same IS in the adaptive sampling formula (Eq. 9) is method construction, not a derivation that makes a predicted quantity identical to an input. The paper does not invoke a self-citation chain: none of the references are authored by the present authors, and no uniqueness claim is imported from the authors' prior work. The most serious technical concern is that Eqs. 9-10 do not, as written, enforce a fixed total keep ratio (floor bounds each bin but cannot make Σq_n equal the target), which threatens the validity of the fixed-ρ comparisons; that is an experimental-control/correctness issue, not a circularity. Because the central accuracy claims are external benchmark comparisons against DM and DQ, not consequences of the definitions, the circularity score is 0.
Assumptions & free parameters
free parameters (6)
- alpha =
not reported per architecture or dataset
- patch size L =
not specified
- temperature tau =
not specified
- gradient operator G =
not specified
- number of bins m =
10 (from DQ)
- drop ratio theta =
25% (from DQ)
assumptions (5)
- domain assumption DQ bin-generation theory: earlier bins are more representative, later bins are more diverse.
- ad hoc to paper Texture level monotonically tracks trajectory closeness to the expert model.
- domain assumption A contrastive instance-discrimination score measures bin diversity.
- ad hoc to paper Normalized RS and DS add with equal weight into a valid importance score.
- domain assumption Validation accuracy is the right criterion for selecting alpha and reporting gains.
Cite this review
Pith. "Pith review of Adaptive Dataset Quantization." pith.science (2026). https://pith.science/paper/2OM7QB5Z
@misc{pith2026241216895,
author = {Pith},
title = {Pith review of: Adaptive Dataset Quantization},
year = {2026},
howpublished = {\url{https://pith.science/paper/2OM7QB5Z}},
note = {Machine review of arXiv:2412.16895}
}
read the original abstract
Contemporary deep learning, characterized by the training of cumbersome neural networks on massive datasets, confronts substantial computational hurdles. To alleviate heavy data storage burdens on limited hardware resources, numerous dataset compression methods such as dataset distillation (DD) and coreset selection have emerged to obtain a compact but informative dataset through synthesis or selection for efficient training. However, DD involves an expensive optimization procedure and exhibits limited generalization across unseen architectures, while coreset selection is limited by its low data keep ratio and reliance on heuristics, hindering its practicality and feasibility. To address these limitations, we introduce a newly versatile framework for dataset compression, namely Adaptive Dataset Quantization (ADQ). Specifically, we first identify the sub-optimal performance of naive Dataset Quantization (DQ), which relies on uniform sampling and overlooks the varying importance of each generated bin. Subsequently, we propose a novel adaptive sampling strategy through the evaluation of generated bins' representativeness score, diversity score and importance score, where the former two scores are quantified by the texture level and contrastive learning-based techniques, respectively. Extensive experiments demonstrate that our method not only exhibits superior generalization capability across different architectures, but also attains state-of-the-art results, surpassing DQ by average 3\% on various datasets.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Agarwal, S.; Arora, H.; Anand, S.; and Arora, C. 2020. Contextual Diversity for Active Learning. In Computer Vision - ECCV 2020 - 16th European Conference ECCV , volume 12361, 137--153. Springer
work page 2020
-
[4]
Cazenavette, G.; Wang, T.; Torralba, A.; Efros, A. A.; and Zhu, J. 2022. Dataset Distillation by Matching Training Trajectories. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR , 10708--10717. IEEE
work page 2022
-
[5]
Cazenavette, G.; Wang, T.; Torralba, A.; Efros, A. A.; and Zhu, J. 2023. Generalizing Dataset Distillation via Deep Generative Prior. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR , 3739--3748. IEEE
work page 2023
-
[6]
Ceccarello, M.; Pietracaprina, A.; and Pucci, G. 2018. Fast Coreset-based Diversity Maximization under Matroid Constraints. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining, WSDM , 81--89. ACM
work page 2018
-
[7]
Coleman, C.; Yeh, C.; Mussmann, S.; Mirzasoleiman, B.; Bailis, P.; Liang, P.; Leskovec, J.; and Zaharia, M. 2020. Selection via Proxy: Efficient Data Selection for Deep Learning. In 8th International Conference on Learning Representations, ICLR . OpenReview.net
work page 2020
-
[8]
Cui, J.; Wang, R.; Si, S.; and Hsieh, C. 2023. Scaling Up Dataset Distillation to ImageNet-1K with Constant Memory. In International Conference on Machine Learning, ICML , volume 202, 6565--6590. PMLR
work page 2023
Show all 40 references
-
[9]
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In 9th International Con...
2021
-
[10]
Du, J.; Jiang, Y.; Tan, V. Y. F.; Zhou, J. T.; and Li, H. 2023. Minimizing the Accumulated Trajectory Error to Improve Dataset Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR , 3749--3758. IEEE
2023
-
[11]
Fang, G.; Song, J.; Wang, X.; Shen, C.; Wang, X.; and Song, M. 2021. Contrastive Model Invertion for Data-Free Knolwedge Distillation. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI , 2374--2380
2021
-
[12]
Feldman, V.; and Zhang, C. 2020. What Neural Networks Memorize and Why: Discovering the Long Tail via Influence Estimation. In A Annual Conference on Neural Information Processing Systems, NeurIPS
2020
-
[13]
Guo, C.; Zhao, B.; and Bai, Y. 2022. DeepCore: A Comprehensive Library for Coreset Selection in Deep Learning. In Database and Expert Systems Applications - 33rd International Conference, DEXA , volume 13426, 181--195. Springer
2022
-
[14]
He, K.; Chen, X.; Xie, S.; Li, Y.; Doll \' a r, P.; and Girshick, R. B. 2022. Masked Autoencoders Are Scalable Vision Learners. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR , 15979--15988. IEEE
2022
-
[15]
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR , 770--778. IEEE Computer Society
2016
-
[16]
Ioffe, S.; and Szegedy, C. 2015. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In Bach, F. R.; and Blei, D. M., eds., Proceedings of the 32nd International Conference on Machine Learning, ICML , volume 37, 448--456
2015
-
[17]
K.; Khargoankar, N.; Bilmes, J
Iyer, R. K.; Khargoankar, N.; Bilmes, J. A.; and Asanani, H. 2021. Submodular combinatorial information measures with applications in machine learning. In Algorithmic Learning Theory, volume 132, 722--754. PMLR
2021
-
[18]
Killamsetty, K.; Sivasubramanian, D.; Ramakrishnan, G.; De, A.; and Iyer, R. K. 2021. GRAD-MATCH: Gradient Matching based Data Subset Selection for Efficient Deep Model Training. In Proceedings of the 38th International Conference on Machine Learning, ICML , volume 139, 5464--...
2021
-
[19]
J.; Yun, S.; Song, H.; Jeong, J.; Ha, J.; and Song, H
Kim, J.; Kim, J.; Oh, S. J.; Yun, S.; Song, H.; Jeong, J.; Ha, J.; and Song, H. O. 2022. Dataset Condensation via Efficient Synthetic-Data Parameterization. In International Conference on Machine Learning, ICML , volume 162, 11102--11118. PMLR
2022
-
[20]
Krizhevsky, A.; Hinton, G.; et al. 2009. Learning multiple layers of features from tiny images. Technical report Citeseer
2009
-
[21]
Le, Y.; and Yang, X. 2015. Tiny imagenet visual recognition challenge. Technical Report, 7(7): 3
2015
-
[22]
Lei, S.; and Tao, D. 2024. A Comprehensive Survey of Dataset Distillation. IEEE Trans. Pattern Anal. Mach. Intell. , 46(1): 17--32
2024
-
[23]
P.; Tang, J.; and Liu, H
Li, J.; Cheng, K.; Wang, S.; Morstatter, F.; Trevino, R. P.; Tang, J.; and Liu, H. 2018. Feature Selection: A Data Perspective. ACM Comput. Surv. , 50(6): 94:1--94:45
2018
-
[24]
Liu, Z.; Hu, H.; Lin, Y.; Yao, Z.; Xie, Z.; Wei, Y.; Ning, J.; Cao, Y.; Zhang, Z.; Dong, L.; Wei, F.; and Guo, B. 2022 a . Swin Transformer V2: Scaling Up Capacity and Resolution. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR , 11999--12009. IEEE
2022
-
[25]
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV , 9992--10002. IEEE
2021
-
[26]
Liu, Z.; Mao, H.; Wu, C.; Feichtenhofer, C.; Darrell, T.; and Xie, S. 2022 b . A ConvNet for the 2020s. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR , 11966--11976. IEEE
2022
-
[27]
Paul, M.; Ganguli, S.; and Dziugaite, G. K. 2021. Deep Learning on a Data Diet: Finding Important Examples Early in Training. In Annual Conference on Neural Information Processing Systems, NeurIPS , 20596--20607
2021
-
[28]
S.; Berg, A
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M. S.; Berg, A. C.; and Fei - Fei, L. 2015. ImageNet Large Scale Visual Recognition Challenge. Int. J. Comput. Vis., 115(3): 211--252
2015
-
[29]
G.; Zhu, M.; Zhmoginov, A.; and Chen, L
Sandler, M.; Howard, A. G.; Zhu, M.; Zhmoginov, A.; and Chen, L. 2018. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR , 4510--4520
2018
-
[30]
T.; Trischler, A.; Bengio, Y.; and Gordon, G
Toneva, M.; Sordoni, A.; des Combes, R. T.; Trischler, A.; Bengio, Y.; and Gordon, G. J. 2019. An Empirical Study of Example Forgetting during Deep Neural Network Learning. In 7th International Conference on Learning Representations, ICLR . OpenReview.net
2019
-
[31]
Wan, Z.; Wang, Z.; Wang, Y.; Wang, Z.; Zhu, H.; and Satoh, S. 2024. Contributing Dimension Structure of Deep Feature for Coreset Selection. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI , 9080--9088. AAAI Press
2024
-
[32]
Wang, K.; Zhao, B.; Peng, X.; Zhu, Z.; Yang, S.; Wang, S.; Huang, G.; Bilen, H.; Wang, X.; and You, Y. 2022. CAFE: Learning to Condense Dataset by Aligning Features. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 12186--12195. IEEE
2022
-
[33]
Wang, T.; Zhu, J.; Torralba, A.; and Efros, A. A. 2018. Dataset Distillation. CoRR, abs/1811.10959
2018 arXiv
-
[34]
Zhang, H.; Li, S.; Wang, P.; Zeng, D.; and Ge, S. 2024. M3D: Dataset Condensation by Minimizing Maximum Mean Discrepancy. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI , 9314--9322. AAAI Press
2024
-
[35]
Zhang, L.; Zhang, J.; Lei, B.; Mukherjee, S.; Pan, X.; Zhao, B.; Ding, C.; Li, Y.; and Xu, D. 2023. Accelerating Dataset Distillation via Model Augmentation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023 , 11950--11959. IEEE
2023
-
[36]
Zhao, B.; and Bilen, H. 2023. Dataset Condensation with Distribution Matching. In IEEE/CVF Winter Conference on Applications of Computer Vision, WACV , 6503--6512. IEEE
2023
-
[37]
R.; and Bilen, H
Zhao, B.; Mopuri, K. R.; and Bilen, H. 2021. Dataset Condensation with Gradient Matching. In 9th International Conference on Learning Representations, ICLR . OpenReview.net
2021
-
[38]
Zhao, G.; Li, G.; Qin, Y.; and Yu, Y. 2023. Improved Distribution Matching for Dataset Condensation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR , 7856--7865. IEEE
2023
-
[39]
Zhou, D.; Wang, K.; Gu, J.; Peng, X.; Lian, D.; Zhang, Y.; You, Y.; and Feng, J. 2023. Dataset Quantization. In IEEE/CVF International Conference on Computer Vision, ICCV , 17159--17170. IEEE
2023
-
[40]
Zhou, Y.; Nezhadarya, E.; and Ba, J. 2022. Dataset Distillation using Neural Feature Regression. In Annual Conference on Neural Information Processing Systems NeurIPS
2022
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.