Pith. sign in

REVIEW 26 cited by

Model compression via distillation and quantization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1802.05668 v1 pith:3B3DJ2DW submitted 2018-02-15 cs.NE cs.LG

classification cs.NEcs.LG
keywords distillationquantizationteachercompressionnetworksquantizedaccuracyadvances
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Deep neural networks (DNNs) continue to make significant advances, solving tasks from image classification to translation or reinforcement learning. One aspect of the field receiving considerable attention is efficiently executing deep models in resource-constrained environments, such as mobile or embedded devices. This paper focuses on this problem, and proposes two new compression methods, which jointly leverage weight quantization and distillation of larger teacher networks into smaller student networks. The first method we propose is called quantized distillation and leverages distillation during the training process, by incorporating distillation loss, expressed with respect to the teacher, into the training of a student network whose weights are quantized to a limited set of levels. The second method, differentiable quantization, optimizes the location of quantization points through stochastic gradient descent, to better fit the behavior of the teacher model. We validate both methods through experiments on convolutional and recurrent architectures. We show that quantized shallow students can reach similar accuracy levels to full-precision teacher models, while providing order of magnitude compression, and inference speedup that is linear in the depth reduction. In sum, our results enable DNNs for resource-constrained environments to leverage architecture and accuracy advances developed on more powerful devices.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ELFuzz: Efficient Input Generation via LLM-driven Synthesis Over Fuzzer Space

    cs.CR 2025-06 conditional novelty 7.0 of 10

    ELFuzz automatically evolves LLM-written input generators for large programs, outperforming grammar-based fuzzers in coverage and bug finding on seven benchmarks.

  2. XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the Edge

    cs.CV 2025-06 unverdicted novelty 6.0 of 10

    XTransfer transfers pre-trained models across sensing modalities with few labeled samples by repairing layer-wise feature misalignment and recombining useful layers, achieving top accuracy and lower resource costs.

  3. Loss-Aware Automatic Selection of Structured Pruning Criteria for Deep Neural Network Acceleration

    cs.CV 2025-06 conditional novelty 6.0 of 10

    LAASP prunes neural networks during training by greedily selecting the best layer and filter-importance criterion at each step using the network's loss on a data subset.

  4. Improving Quantization-aware Training of Low-Precision Network via Block Replacement on Full-Precision Counterpart

    cs.LG 2024-12 conditional novelty 6.0 of 10

    BWRF improves quantization-aware training by grafting full-precision blocks onto the low-precision model during training, producing mixed-precision guides that raise ImageNet and CIFAR-10 accuracy at 2 to 4 bits.

  5. DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization

    cs.CV 2024-12 unverdicted novelty 6.0 of 10

    DOLLAR combines variational score and consistency distillation for few-step video generation plus latent reward optimization, reporting 82.57 VBench score and up to 278x speedup over the teacher diffusion model for 12...

  6. RBCN: Rectified Binary Convolutional Networks for Enhancing the Performance of 1-bit DCNNs

    cs.CV 2019-08 conditional novelty 6.0 of 10

    A GAN-guided rectified training scheme narrows the accuracy gap between binary and full-precision convolutional networks on classification and tracking.

  7. Fast Tensorization of Neural Networks via Slice-wise Feature Distillation

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    A slice-wise feature distillation framework for independent tensorization of neural network slices to achieve scalable compression with reduced fine-tuning costs.

  8. Low-Rank Augmented Implicit Neural Representation for Unsupervised High-Dimensional Quantitative MRI Reconstruction

    eess.IV 2025-06 conditional novelty 5.0 of 10

    LoREIN couples implicit neural representations with low-rank temporal subspace modeling to reconstruct quantitative MRI maps and weighted images directly from undersampled k-space.

  9. ReverB-SNN: Reversing Bit of the Weight and Activation for Spiking Neural Networks

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ReverB-SNN replaces binary spikes with real-valued spikes and real weights with binary weights, keeping SNN inference addition-only while improving accuracy.

  10. EfficientLLM: Efficiency in Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A large-scale benchmark of LLM efficiency techniques finds that every method trades off one resource for another, with the best choice depending on model scale, task, and hardware.

  11. Adaptive Pruning of Pretrained Transformer via Differential Inclusions

    cs.LG 2025-01 conditional novelty 5.0 of 10

    A single differential-inclusion search over masks produces a whole family of pruned transformers at different sparsity levels from one pretrained model.

  12. Optimising TinyML with Quantization and Distillation of Transformer and Mamba Models for Indoor Localisation on Edge Devices

    cs.LG 2024-12 conditional novelty 5.0 of 10

    A quantized transformer and a compact Mamba model can classify indoor location with moderate accuracy within 32-64 KB model sizes, but on-device RAM usage is not measured.

  13. Patient Knowledge Distillation for BERT Model Compression

    cs.CL 2019-08 conditional novelty 5.0 of 10

    Distilling BERT through several intermediate hidden layers (Patient-KD) improves a shallow student's accuracy on GLUE and RACE compared with last-layer-only distillation.

  14. Knowledge Distillation for Sensing-Assisted Long-Term Beam Tracking in mmWave Communications

    eess.SP 2025-09 unverdicted novelty 4.0 of 10

    Knowledge distillation creates a compact neural network for long-term beam tracking in mmWave communications that matches a larger teacher's accuracy with far fewer parameters and shorter input sequences.

  15. Resource-Efficient Automatic Software Vulnerability Assessment via Knowledge Distillation and Particle Swarm Optimization

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A PSO-guided knowledge distillation framework compresses a CodeBERT vulnerability assessor to 0.6% of its original size while retaining 89.3% of its accuracy.

  16. The Promise of Spiking Neural Networks for Ubiquitous Computing: A Survey and New Perspectives

    cs.NE 2025-06 conditional novelty 4.0 of 10

    A survey of 76 spiking-neural-network papers on time-series sensor data, organized into six application domains, with recommendations for software and neuromorphic hardware.

  17. Tensorization is a powerful but underexplored tool for compression and interpretability of neural networks

    cs.LG 2025-05 conditional novelty 4.0 of 10

    The paper makes the case that tensorized neural networks offer valuable compression, scaling, and interpretability advantages that the deep learning community has not yet fully exploited.

  18. On Hardening DNNs against Noisy Computations

    cs.LG 2025-01 conditional novelty 4.0 of 10

    On CIFAR-10, quantization-aware training with large constant scaling factors improves noise robustness, but noisy training (injecting matching Gaussian noise during training) gives far larger robustness gains, and qua...

  19. Pan-infection Foundation Framework Enables Multiple Pathogen Prediction

    cs.LG 2024-12 reject novelty 4.0 of 10

    A teacher-student knowledge distillation framework trained on 11,247 blood transcriptomes reports high AUCs for pan-infection, four pathogens, and sepsis diagnosis.

  20. Multi-Modality Distillation via Learning the teacher's modality-level Gram Matrix

    cs.AI 2021-12 unverdicted novelty 4.0 of 10

    Proposes a modality relation distillation method that transfers teacher modality relationships via the modality-level Gram Matrix.

  21. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

  22. Recent Advances and Trends in Learning-based 3D Representations

    cs.CV 2026-06 unverdicted novelty 3.0 of 10

    A survey of 3D representation families that highlights the shift from discrete explicit formats to continuous implicit neural and primitive-based ones.

  23. Vision Generalist Model: A Survey

    cs.CV 2025-06 conditional novelty 3.0 of 10

    A structured review of vision generalist models, classifying them into encoding-based and sequence-to-sequence frameworks and summarizing datasets, benchmarks, techniques, and open problems.

  24. A Survey on Foundation Models for Personalized Federated Intelligence

    cs.AI 2025-05 unverdicted novelty 3.0 of 10

    The survey introduces personalized federated intelligence (PFI) as a framework integrating federated learning and foundation models to support privacy-aware personalization of AI models.

  25. Frugal Machine Learning for Energy-efficient, and Resource-aware Artificial Intelligence

    cs.LG 2025-06 conditional novelty 1.0 of 10

    A survey paper that defines and categorizes Frugal Machine Learning methods but introduces no new techniques or empirical results.

  26. Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques

    cs.LG 2025-05 conditional novelty 1.0 of 10

    A survey of knowledge distillation, quantization, and pruning for compressing LLMs to run on resource-constrained edge devices, with no new experimental results.

Pith tools