Pith. sign in

REVIEW 1 cited by

Dynamic Network Surgery for Efficient DNNs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1608.04493 v2 pith:VKZPJGQX submitted 2016-08-16 cs.NE cs.CVcs.LG

classification cs.NEcs.CVcs.LG
keywords networkmethodpruningconnectiondeepdynamicmakingmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Deep learning has become a ubiquitous technology to improve machine intelligence. However, most of the existing deep models are structurally very complex, making them difficult to be deployed on the mobile platforms with limited computational power. In this paper, we propose a novel network compression method called dynamic network surgery, which can remarkably reduce the network complexity by making on-the-fly connection pruning. Unlike the previous methods which accomplish this task in a greedy way, we properly incorporate connection splicing into the whole process to avoid incorrect pruning and make it as a continual network maintenance. The effectiveness of our method is proved with experiments. Without any accuracy loss, our method can efficiently compress the number of parameters in LeNet-5 and AlexNet by a factor of $\bm{108}\times$ and $\bm{17.7}\times$ respectively, proving that it outperforms the recent pruning method by considerable margins. Code and some models are available at https://github.com/yiwenguo/Dynamic-Network-Surgery.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Compression of Language Models for Code: An Empirical Study on CodeBERT

    cs.SE 2024-12 conditional novelty 5.0 of 10

    On CodeBERT, quantization best preserves effectiveness while cutting size, distillation best improves latency, and pruning only pays off in specific CPU configurations.

Pith tools