REVIEW 3 major objections 8 minor 77 references
A 1Mb mixed-precision quantized encoder for image classification and patch-based compression
T0 review · 3 major / 8 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A reconfigurable 1 Mb mixed-precision encoder doubles as a CIFAR-10 classifier and a 0.25-bpp image compressor.
desk verdict A credible mixed-precision encoder with a useful compression twist, but the headline 1Mb figure is weights-only and is compared to full on-chip memory of fabricated chips. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Nonlinear Quantized Encoder (NQE): a VGG-style 7-layer network in which the first convolutional layers use quinary (5-level) weights, the middle layers use ternary weights, and the rest use binary weights, with activations at 1 or 2 bits. Three mechanisms make it work: (1) histogram-equidistributed symmetric linear quantization, which re-estimates the step size each epoch so quantized weight levels are roughly equally populated rather than assuming a Gaussian weight distribution; (2) Half-Wave Most-Significant-Bit activation, which turns a value into one of four outputs {0, 1/3, 2/3, 1} using only 2 bits; and (3) layer-shared Bit-Shift Normalization, which replaces every Batch Normalization affine transform with a single power-of-2 bitshift per layer, so most normalization disappears at inference time. Structural group-wise convolution on the last convolutional layer and a depthwise-convolution-plus-small-FC bottleneck cut the parameter count. For reconstruction, the decoder PURENET first upsamples each patch's code independently, then aggregates patches and refines the full frame, which is what removes block artifacts.
What would settle it
Take the published NQE topology at $F=64$ and add the activation feature maps, patch input/output buffers, and any weight-decode storage needed for an ASIC; if the resulting memory budget clearly exceeds 1 Mb, the '1Mb encoder' claim is false at the system level. A lighter check is to retrain CIFAR-10 without the per-epoch histogram update of tau and see whether accuracy drops by more than the reported gain over the binary baseline.
Extended reading notes
Core claim
The central claim is that a single hand-crafted, mixed-precision encoder topology can be reused across tasks when its weights are retrained. On CIFAR-10, with the feature-scale hyperparameter $F=64$, the Nonlinear Quantized Encoder (NQE) plus a one-layer classifier reaches 87.48% average accuracy while storing about 1.073 Mb of weights using naive encoding; this is 5% above a fully binarized version of the same topology and, the paper reports, better than several fabricated accelerators at comparable or smaller memory. For compression, the same NQE encodes non-overlapping 32-by-32 patches as 256-bit binary vectors, and the proposed PURENET decoder stitches patch codes into a full frame, reaching 20.76 dB PSNR and 0.8136 MS-SSIM on VGA DIV2K at 0.25 bpp, ahead of JPEG and of block-based compressed sensing with WD-TV3D reconstruction. The contribution is not a new network family but a set of hardware-friendly algorithmic pieces: histogram-equidistributed quantization for quinary and ternary weights, a 2-bit Half-Wave MSB activation, layer-shared Bit-Shift Normalization replacing Batch Normalization, group-wise convolution pruning, and a depthwise-convolution alternative to the dense bottleneck, which together make one small encoder viable for both semantic and pixel-level tasks.
Load-bearing premise
The 1 Mb figure counts only the encoder's weight memory with naive 3/2/1-bit encoding; no synthesis, activation-memory accounting, or power and area estimate is given, so the claim that the whole encoder requires only 1 Mb as a chip is not yet demonstrated.
Editorial extensions
If this is right
- A classifier and a compressor can share one reconfigurable encoder chip; switching tasks only means loading different weights.
- At 0.25 bpp, the fully quantized encoder outperforms JPEG and block-based compressed sensing on VGA DIV2K in both PSNR and MS-SSIM, so a very simple near-sensor encoder can produce usable images.
- The layer-shared Bit-Shift Normalization removes the per-channel affine transforms of Batch Normalization, simplifying future ASIC implementations while keeping accuracy loss small.
- Entropy-coding quinary and ternary weights, rather than naive 3-bit/2-bit storage, could push the same model below 1 Mb of weight memory.
- Because the encoder is patch-wise, compression memory and compute scale with patch count rather than full-image resolution.
Reading between the lines
- The 1 Mb figure would face a harder test in a full ASIC memory budget; the paper counts only weight memory with naive encoding, so activation storage, patch buffers, and I/O could lift the real on-chip number substantially.
- The same reconfigurable encoder idea likely extends to other small image tasks such as detection or denoising, because the encoder weights are trained separately per task; the paper only demonstrates classification and compression.
- The entropy-coding remark suggests a fast follow-up: measure the actual bitstream entropy of the quantized weights and compare it with the 1.073 Mb naive figure to see how far below 1 Mb a practical chip could go.
- PURENET's full-frame refinement is what removes block artifacts, so future ultra-low-precision encoders could become even simpler if the decoder is given more context.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a mixed-precision quantized encoder (NQE) with quinary/ternary/binary weights and 1-2 bit activations, aimed at a hardware-efficient ASIC design. The encoder uses three algorithmic enablers: histogram-equidistributed quantization that adaptively sets the quantizer step, a half-wave MSB activation for 2-bit activations, and a layer-shared bit-shift normalization (BSN) that replaces batch normalization. The same encoder topology is used for CIFAR-10 classification (with a classifier head) and for patch-based image compression (with a remote full-frame decoder called PURENET). For a configuration with F=64, the weight memory is reported as about 1 Mb, the CIFAR-10 accuracy is 87.48% (average over 3 runs), and at 0.25 bpp the compression pipeline achieves 20.76 dB PSNR and 0.8136 MS-SSIM on VGA DIV2K, outperforming JPEG and the compressed-sensing baseline WD-TV3D. The paper also presents memory-accuracy trade-off curves for different feature-map scales and a comparison of bottleneck alternatives in appendices.
Significance. If the results hold, the paper demonstrates a single reconfigurable low-precision encoder topology that serves two distinct tasks with very low weight memory, using hardware-friendly operations (bitshift, MSB extraction, group convolution). The patch-based compression use case for a quantized encoder is relatively novel, and the reported compression results against JPEG and BCS baselines are concrete and reproducible in principle. The paper is also transparent about the weight-only nature of the memory figure in the main text, and it provides useful design-space exploration in Appendices A and B. However, the significance is currently tempered by two gaps: the headline '1Mb' is not a demonstrated system-level on-chip memory figure, and the two main algorithmic contributions (histogram-equalized quantization and BSN) are not isolated by ablations, so their individual impact on the reported accuracy is unverified.
major comments (3)
- [V.A, Table III] The headline '1Mb' figure is computed as the sum of the weight tensors under naive per-layer bitwidths (stated in Section V as 'with naive weights encoding' and detailed in Appendix A/Table V), yet Table III reports it in a column labeled 'On-chip Memory (Mb)' and compares it directly with fabricated accelerators whose on-chip memories include activation storage, data buffers, and other system-level overheads. Since the abstract and title claim that the encoder 'only requires 1Mb', the system-level memory claim is not demonstrated: there is no synthesis, no activation-memory accounting, and no estimate of line/column buffers, the 8-bit input plane, the 2-bit feature maps, HWMSB reference shifts, or BSN constants. The paper should either re-scope the claim to 'weight memory' throughout the title, abstract, and Table III, or provide a complete on-chip memory estimate for a concrete accelerator dataflow.
- [III.A, III.C, V.A] The two main algorithmic contributions—histogram-equidistributed quantization and layer-shared BSN—are not isolated by any ablation. There is no experiment comparing the proposed adaptive Delta against a fixed Delta (e.g., the TWN 0.7 rule) or against another adaptation scheme, and no comparison of BSN against standard BN or a fused-BN baseline in the final accuracy. Section II.A sets a target of under 1% degradation for the BN replacement, but the paper never reports the accuracy with standard BN before the BSN swap. Without these ablations, the individual contributions of the proposed quantizer and BSN to the 87.5% result are not established, and the claim that these enablers 'stabilize' training and 'simplify' hardware is supported only by qualitative arguments.
- [V.B, Table IV] The compression comparison is presented as outperforming 'patch-based state-of-the-art techniques', but the only patch-based methods beaten are JPEG and WD-TV3D; the stronger learning-based methods (RCAE, CAEM-PSNR, CAEM-MS-SSIM) outperform the proposed approach and are full-resolution, non-quantized systems. The paper acknowledges this, which is fair, but the conclusion that the proposed method achieves a 'good quality of service versus computational complexity' is not quantified: no decoder-side complexity (MACs, parameters, or runtime) is reported for PURENET, and no encoder-side complexity comparison is made against JPEG/JPEG2000 beyond a qualitative discussion. A quantitative complexity-accuracy trade-off plot, or at least decoder MAC/parameter counts, would make this claim testable.
minor comments (8)
- [III.B, Eq. (3)] Equation (3) states the condition 'if x >= 1/8', but the function is supposed to map into [-1,1] for all inputs; for large negative x the 'otherwise' branch 8x/3 produces values far below -1. The condition should presumably be '|x| >= 1/8', consistent with Eq. (4). Please correct the definition.
- [III.C] The text says the 0.9-quantile of the BN scales is chosen 'at each layer' and then refers to 'the unique scale for all the BSNs' and a 'single BSN transform'. It should be clarified whether the bitshift scale is per-layer or truly global; Table III lists the share level as 'layer', but the wording is ambiguous.
- [Table III] The table column alignment is difficult to read, especially for the 'Batch Norm', 'Bias-precision', and 'Implementation' rows; it appears that entries may be shifted across columns. Please reformat so each column is unambiguous.
- [V.A] The paper reports average accuracies over 3 realizations but gives no standard deviation or per-run values. Since several comparisons in Table V involve differences of about 0.2% (e.g., 87.48% vs. 87.69%), reporting error bars or individual runs would substantially strengthen the quantitative claims.
- [V.B, Table IV] Please state explicitly whether the 0.25 bpp for the NQE-PURENET pipeline is the raw, non-entropy-coded bitrate, while the JPEG/JPEG2000 bitrates are after entropy coding. This is important for a fair interpretation of the comparison.
- [III.A] In the approximation for the ternary quantization step, the formula appears as 'Delta = |q1|+q2 2'; the missing division symbol makes it ambiguous. It should read (|q1|+q2)/2.
- [References] Reference [70] (Kingma and Ba, Adam) should be dated 2015 or 2017 depending on the cited version; the current citation year is unconventional for this well-known work.
- [II.A] There is a typo in the phrase 'improve the harware/algorithm trade-offs'; 'harware' should be 'hardware'.
Circularity Check
No significant circularity: the paper's central results are measured outcomes against external benchmarks, and the adaptive quantizer is tuned to weight statistics rather than to the target accuracy metrics.
full rationale
The paper's main claims are empirical: the mixed-precision NQE reaches 87.48% on CIFAR-10 with about 1 Mb of weight memory, and the NQE-PURENET system reaches 20.76 dB PSNR at 0.25 bpp on DIV2K. These are measured results, not derived from definitions. The 1 Mb figure is the sum of per-layer weight bits, consistent with Table I and Appendix A, and the paper explicitly labels it as 'naive weights encoding' in Section V; comparing this weight-memory figure to the on-chip memory of fabricated accelerators is a scope mismatch but not a circular step. The quantization step Delta is estimated from layer-wise weight histograms via Equation (2), and the scaling factor tau is updated from proxy weights during training; no target accuracy is used to set the quantizer, so there is no fitted-input-called-prediction pattern. The BSN is obtained by first training with standard BN, then approximating the scale by a bitshift and retraining; this is an implementation simplification rather than a circular derivation. Compression benchmarks are external: JPEG, JPEG2000, WD-TV3D, RCAE, and CAEM are independent baselines evaluated on the same DIV2K test set. Self-citations [46] and [52] appear only as related-work context or as inspiration for a baseline, not as load-bearing justification of the central claims. There is no uniqueness theorem, no ansatz smuggled in via self-citation, and no renaming of a known result as a new derivation. Therefore the derivation chain is self-contained and no circularity is present.
Assumptions & free parameters
free parameters (5)
- BSN 0.9-quantile =
0.9
- HWMSB integer bias =
4
- HWMSB normalization factor =
3
- Number of groups in group convolution =
4
- Feature map scale F =
64
assumptions (4)
- domain assumption Weights distributions are approximately symmetric around zero, so quantile points can be mirrored.
- standard math Straight-through estimator provides a usable gradient proxy for non-differentiable quantizers.
- domain assumption Batch Normalization is required for quantized networks to converge well.
- ad hoc to paper The 0.9-quantile of BN scale factors is a good single scale for all layers.
Cite this review
Pith. "Pith review of A 1Mb mixed-precision quantized encoder for image classification and patch-based compression." pith.science (2026). https://pith.science/paper/2JMNDK7P
@misc{pith2026250105097,
author = {Pith},
title = {Pith review of: A 1Mb mixed-precision quantized encoder for image classification and patch-based compression},
year = {2026},
howpublished = {\url{https://pith.science/paper/2JMNDK7P}},
note = {Machine review of arXiv:2501.05097}
}
read the original abstract
Even if Application-Specific Integrated Circuits (ASIC) have proven to be a relevant choice for integrating inference at the edge, they are often limited in terms of applicability. In this paper, we demonstrate that an ASIC neural network accelerator dedicated to image processing can be applied to multiple tasks of different levels: image classification and compression, while requiring a very limited hardware. The key component is a reconfigurable, mixed-precision (3b/2b/1b) encoder that takes advantage of proper weight and activation quantizations combined with convolutional layer structural pruning to lower hardware-related constraints (memory and computing). We introduce an automatic adaptation of linear symmetric quantizer scaling factors to perform quantized levels equalization, aiming at stabilizing quinary and ternary weights training. In addition, a proposed layer-shared Bit-Shift Normalization significantly simplifies the implementation of the hardware-expensive Batch Normalization. For a specific configuration in which the encoder design only requires 1Mb, the classification accuracy reaches 87.5% on CIFAR-10. Besides, we also show that this quantized encoder can be used to compress image patch-by-patch while the reconstruction can performed remotely, by a dedicated full-frame decoder. This solution typically enables an end-to-end compression almost without any block artifacts, outperforming patch-based state-of-the-art techniques employing a patch-constant bitrate.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Efficient Processing of Deep Neural Networks: A Tutorial and Survey,
V . Sze, Y .-H. Chen, T.-J. Yang, and J. S. Emer, “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Proceedings of the IEEE, vol. 105, no. 12, pp. 2295–2329, Dec. 2017
work page 2017
-
[2]
Fast Training of Convolutional Networks through FFTs,
M. Mathieu, M. Henaff, and Y . LeCun, “Fast Training of Convolutional Networks through FFTs,” International Conference on Learning Repre- sentation (ICLR), 2014
work page 2014
-
[3]
Minimizing computation in convolutional neural networks,
J. Cong and B. Xiao, “Minimizing computation in convolutional neural networks,” in Artificial Neural Networks and Machine Learn- ing – ICANN 2014 , S. Wermter, C. Weber, W. Duch, T. Honkela, P. Koprinkova-Hristova, S. Magg, G. Palm, and A. E. P. Villa, Eds. Cham: Springer International Publishing, 2014, pp. 281–290
work page 2014
-
[4]
Fast Algorithms for Convolutional Neural Networks,
A. Lavin and S. Gray, “Fast Algorithms for Convolutional Neural Networks,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2016, pp. 4013–4021
work page 2016
-
[5]
Deep Convolu- tional Neural Network Architecture With Reconfigurable Computation Patterns,
F. Tu, S. Yin, P. Ouyang, S. Tang, L. Liu, and S. Wei, “Deep Convolu- tional Neural Network Architecture With Reconfigurable Computation Patterns,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 25, no. 8, pp. 2220–2233, Aug. 2017
work page 2017
-
[6]
Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,
Y .-H. Chen, J. Emer, and V . Sze, “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), Jun. 2016, pp. 367–379, iSSN: 1063-6897
work page 2016
-
[7]
A. Aimar, H. Mostafa, E. Calabrese, A. Rios-Navarro, R. Tapiador- Morales, I.-A. Lungu, M. B. Milde, F. Corradi, A. Linares-Barranco, S.-C. Liu, and T. Delbruck, “NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps,” IEEE Transactions on Neural Networks and Learning Systems , vol. 30, no. 3, pp. 644...
work page 2019
-
[8]
UNPU: An Energy-Efficient Deep Neural Network Accelerator With Fully Variable Weight Bit Precision,
J. Lee, C. Kim, S. Kang, D. Shin, S. Kim, and H.-J. Yoo, “UNPU: An Energy-Efficient Deep Neural Network Accelerator With Fully Variable Weight Bit Precision,” IEEE Journal of Solid-State Circuits , vol. 54, no. 1, pp. 173–185, Jan. 2019
work page 2019
Show all 77 references
-
[9]
15.2 A 2.75-to-75.9TOPS/W Computing- in-Memory NN Processor Supporting Set-Associate Block-Wise Zero Skipping and Ping-Pong CIM with Simultaneous Computation and Weight Updating,
J. Yue, X. Feng, Y . He, Y . Huang, Y . Wang, Z. Yuan, M. Zhan, J. Liu, J.-W. Su, Y .-L. Chung, P.-C. Wu, L.-Y . Hung, M.-F. Chang, N. Sun, X. Li, H. Yang, and Y . Liu, “15.2 A 2.75-to-75.9TOPS/W Computing- in-Memory NN Processor Supporting Set-Associate Block-Wise Zero Skippi...
2021
-
[10]
A Resource-Efficient Inference Accelerator for Binary Convolutional Neural Networks,
T.-H. Kim and J. Shin, “A Resource-Efficient Inference Accelerator for Binary Convolutional Neural Networks,” IEEE Transactions on Circuits and Systems II: Express Briefs , vol. 68, no. 1, pp. 451–455, Jan. 2021, conference Name: IEEE Transactions on Circuits and Systems II: E...
2021
-
[11]
YodaNN: An Ultra- Low Power Convolutional Neural Network Accelerator Based on Binary Weights,
R. Andri, L. Cavigelli, D. Rossi, and L. Benini, “YodaNN: An Ultra- Low Power Convolutional Neural Network Accelerator Based on Binary Weights,” in 2016 IEEE Computer Society Annual Symposium on VLSI , Jul. 2016, pp. 236–241
2016
-
[12]
An Ultra-Low-Power Analog-Digital Hybrid CNN Face Recognition Processor Integrated with a CIS for Always-on Mobile Devices,
J.-H. Kim, C. Kim, K. Kim, and H.-J. Yoo, “An Ultra-Low-Power Analog-Digital Hybrid CNN Face Recognition Processor Integrated with a CIS for Always-on Mobile Devices,” in 2019 IEEE International Symposium on Circuits and Systems (ISCAS) , May 2019, pp. 1–5
2019
-
[13]
A 3.0µW@5fps QQVGA Self-Controlled Wake-Up Imager with On-Chip Motion Detection, Auto-Exposure and Object Recognition,
A. Verdant, W. Guicquero, N. Royer, G. Moritz, S. Martin, F. Lepin, S. Choisnet, F. Guellec, B. Deschamps, S. Clerc, and J. Chossat, “A 3.0µW@5fps QQVGA Self-Controlled Wake-Up Imager with On-Chip Motion Detection, Auto-Exposure and Object Recognition,” in 2020 IEEE Symposium ...
2020
-
[14]
Lefebvre, L
M. Lefebvre, L. Moreau, R. Dekimpe, and D. Bol, “7.7 A 0.2-to- 3.6TOPS/W Programmable Convolutional Imager SoC with In-Sensor Current-Domain Ternary-Weighted MAC Operations for Feature Ex- traction and Region-of-Interest Detection,” in 2021 IEEE International Solid- State Circ...
2021
-
[15]
A multi- task hardwired accelerator for face detection and alignment,
H. Mo, L. Liu, W. Zhu, Q. Li, H. Liu, S. Yin, and S. Wei, “A multi- task hardwired accelerator for face detection and alignment,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 30, no. 11, pp. 4284–4298, 2020
2020
-
[16]
A 617 TOPS/W All Digital Binary Neural Network Accelerator in 10nm FinFET CMOS,
P. C. Knag, G. K. Chen, H. E. Sumbul, R. Kumar, M. A. Anders, H. Kaul, S. K. Hsu, A. Agarwal, M. Kar, S. Kim, and R. K. Krishna- murthy, “A 617 TOPS/W All Digital Binary Neural Network Accelerator in 10nm FinFET CMOS,” in 2020 IEEE Symposium on VLSI Circuits , Jun. 2020, pp. 1...
2020
-
[17]
A 68-mw 2.2 tops/w low bit width and multiplierless dcnn object detection processor for visually impaired peo- ple,
X. Chen, J. Xu, and Z. Yu, “A 68-mw 2.2 tops/w low bit width and multiplierless dcnn object detection processor for visually impaired peo- ple,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 29, no. 11, pp. 3444–3453, 2019
2019
-
[18]
Layer-specific optimization for mixed data flow with mixed precision in fpga design for cnn-based object detectors,
D. T. Nguyen, H. Kim, and H.-J. Lee, “Layer-specific optimization for mixed data flow with mixed precision in fpga design for cnn-based object detectors,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 6, pp. 2450–2464, 2021
2021
-
[19]
Neural network accelerator comparison,
K. Guo, W. Li, K. Zhong, Z. Zhu, S. Zeng, S. Han, Y . Xie, P. Debacker, M. Verhelst, and Y . Wang, “Neural network accelerator comparison,” Tech. Rep. [Online]. Available: https://nicsefc.ee.tsinghua. edu.cn/projects/neural-network-accelerator/
-
[20]
Convolutional neu- ral networks using logarithmic data representation,
D. Miyashita, E. G. Lee, and B. Murmann, “Convolutional neu- ral networks using logarithmic data representation,” ArXiv, vol. abs/1603.01025, 2016
2016 arXiv
-
[21]
Minimum energy quantized neural networks,
B. Moons, K. Goetschalckx, N. Van Berckelaer, and M. Verhelst, “Minimum energy quantized neural networks,” in 2017 51st Asilomar Conference on Signals, Systems, and Computers , Oct. 2017
2017
-
[22]
Binarized Neural Networks,
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y . Bengio, “Binarized Neural Networks,” in Advances in Neural Information Pro- cessing Systems 29, D. D. Lee, M. Sugiyama, U. V . Luxburg, I. Guyon, and R. Garnett, Eds. Curran Associates, Inc., 2016, pp. 4107–4115
2016
-
[23]
XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,
M. Rastegari, V . Ordonez, J. Redmon, and A. Farhadi, “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” arXiv:1603.05279 [cs], Aug. 2016
2016 arXiv
-
[24]
Ternary Weight Networks,
F. Li, B. Zhang, and B. Liu, “Ternary Weight Networks,” arXiv:1605.04711 [cs], Nov. 2016
2016 arXiv
-
[25]
Trained Ternary Quantiza- tion,
C. Zhu, S. Han, H. Mao, and W. J. Dally, “Trained Ternary Quantiza- tion,” arXiv:1612.01064 [cs], Feb. 2017
2017 arXiv
-
[26]
Mixed Precision DNNs: All you need is a good parametrization,
S. Uhlich, L. Mauch, F. Cardinaux, K. Yoshiyama, J. A. Garcia, S. Tiedemann, T. Kemp, and A. Nakamura, “Mixed Precision DNNs: All you need is a good parametrization,” arXiv:1905.11452 [cs, stat] , May 2020
1905 arXiv
-
[27]
Relaxed Quantization for Discretized NeuralL Networks,
C. Louizos, M. Reisser, T. Blankevoort, E. Gavves, and M. Welling, “Relaxed Quantization for Discretized NeuralL Networks,” p. 15, 2019
2019
-
[28]
Linear symmetric quantization of neural networks for low-precision integer hardware,
X. Zhao, Y . Wang, X. Cai, C. Liu, and L. Zhang, “Linear symmetric quantization of neural networks for low-precision integer hardware,” p. 16, 2020
2020
-
[29]
Learned step size quantization,
S. K. Esser, J. McKinstry, D. Bablani, R. Appuswamy, and D. Modha, “Learned step size quantization,” ICLR, 2020
2020
-
[30]
Batch Normalization: Accelerating Deep Net- work Training by Reducing Internal Covariate Shift,
S. Ioffe and C. Szegedy, “Batch Normalization: Accelerating Deep Net- work Training by Reducing Internal Covariate Shift,” arXiv:1502.03167 [cs], Mar. 2015
2015 arXiv
-
[31]
An Always-On 3.8 µ J /86% CIFAR-10 Mixed-Signal Binary CNN Processor With All Memory on Chip in 28-nm CMOS,
D. Bankman, L. Yang, B. Moons, M. Verhelst, and B. Murmann, “An Always-On 3.8 µ J /86% CIFAR-10 Mixed-Signal Binary CNN Processor With All Memory on Chip in 28-nm CMOS,” IEEE Journal of Solid-State Circuits , vol. 54, no. 1, pp. 158–172, Jan. 2019
2019
-
[32]
Riptide: Fast end-to-end binarized neural networks,
J. Fromm, M. Cowan, M. Philipose, L. Ceze, and S. Patel, “Riptide: Fast end-to-end binarized neural networks,” in Proceedings of Machine Learning and Systems , 2020
2020
-
[33]
A 64-Tile 2.4-Mb In-Memory-Computing CNN Accelerator Employing Charge-Domain Compute,
H. Valavi, P. J. Ramadge, E. Nestler, and N. Verma, “A 64-Tile 2.4-Mb In-Memory-Computing CNN Accelerator Employing Charge-Domain Compute,” IEEE Journal of Solid-State Circuits , vol. 54, no. 6, pp. 1789–1799, Jun. 2019, conference Name: IEEE Journal of Solid-State Circuits
2019
-
[34]
A programmable heterogeneous microprocessor based on bit-scalable in-memory comput- ing,
H. Jia, H. Valavi, Y . Tang, J. Zhang, and N. Verma, “A programmable heterogeneous microprocessor based on bit-scalable in-memory comput- ing,” IEEE Journal of Solid-State Circuits, vol. 55, no. 9, pp. 2609–2621, 2020
2020
-
[35]
Learning both Weights and Connections for Efficient Neural Network,
S. Han, J. Pool, J. Tran, and W. Dally, “Learning both Weights and Connections for Efficient Neural Network,” in Advances in Neural Information Processing Systems 28 , C. Cortes, N. D. Lawrence, D. D. Lee, M. Sugiyama, and R. Garnett, Eds. Curran Associates, Inc., 2015, pp. 1135–1143
2015
-
[36]
Channel Pruning for Accelerating Very Deep Neural Networks,
Y . He, X. Zhang, and J. Sun, “Channel Pruning for Accelerating Very Deep Neural Networks,” in 2017 IEEE International Conference on Computer Vision (ICCV) . Venice: IEEE, Oct. 2017, pp. 1398–1406
2017
-
[37]
Learning Structured Sparsity in Deep Neural Networks,
W. Wen, C. Wu, Y . Wang, Y . Chen, and H. Li, “Learning Structured Sparsity in Deep Neural Networks,” in Advances in Neural Information Processing Systems 29 , D. D. Lee, M. Sugiyama, U. V . Luxburg, I. Guyon, and R. Garnett, Eds. Curran Associates, Inc., 2016, pp. 2074–2082
2016
-
[38]
Structured Pruning of Neural Networks With Budget-Aware Regularization,
C. Lemaire, A. Achkar, and P.-M. Jodoin, “Structured Pruning of Neural Networks With Budget-Aware Regularization,” in 2019 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR) . Long Beach, CA, USA: IEEE, Jun. 2019, pp. 9100–9108. 14
2019
-
[39]
Accelerator-aware pruning for convolutional neural net- works,
H.-J. Kang, “Accelerator-aware pruning for convolutional neural net- works,” IEEE Transactions on Circuits and Systems for Video Technol- ogy, vol. 30, no. 7, pp. 2093–2103, 2020
2020
-
[40]
Recovering from Random Pruning: On the Plasticity of Deep Convolutional Neural Networks,
D. Mittal, S. Bhardwaj, M. M. Khapra, and B. Ravindran, “Recovering from Random Pruning: On the Plasticity of Deep Convolutional Neural Networks,” in 2018 IEEE Winter Conference on Applications of Com- puter Vision (WACV), Mar. 2018, pp. 848–857
2018
-
[42]
ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices,
X. Zhang, X. Zhou, M. Lin, and J. Sun, “ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Salt Lake City, UT: IEEE, Jun. 2018, pp. 6848–6856
2018
-
[43]
Network in network,
M. Lin, Q. Chen, and S. Yan, “Network in network,” CoRR, vol. abs/1312.4400, 2014
2014 arXiv
-
[44]
Channelnets: Compact and efficient convolutional neural networks via channel-wise convolutions,
H. Gao, Z. Wang, and S. Ji, “Channelnets: Compact and efficient convolutional neural networks via channel-wise convolutions,” IEEE transactions on pattern analysis and machine intelligence , 2018
2018
-
[45]
Fix your classifier: the marginal value of training the last weight layer,
E. Hoffer, I. Hubara, and D. Soudry, “Fix your classifier: the marginal value of training the last weight layer,” ICLR, 2018
2018
-
[46]
Hardware- Friendly Compressive Imaging Based on Random Modulations Per- mutations for Image Acquisition and Classification,
W. Benjilali, W. Guicquero, L. Jacques, and G. Sicard, “Hardware- Friendly Compressive Imaging Based on Random Modulations Per- mutations for Image Acquisition and Classification,” in 2019 IEEE International Conference on Image Processing (ICIP) , Sep. 2019, pp. 2085–2089, iSS...
2019
-
[47]
The JPEG still picture compression standard,
G. Wallace, “The JPEG still picture compression standard,” IEEE Transactions on Consumer Electronics , vol. 38, no. 1, pp. xviii–xxxiv, 1992
1992
-
[48]
J. E. Fowler, S. Mun, and E. W. Tramel, Block-Based Compressed Sensing of Images and Video , 2012
2012
-
[49]
CMOS Image Sensor With Per-Column σδ ADC and Programmable Compressed Sensing,
Y . Oike and A. El Gamal, “CMOS Image Sensor With Per-Column σδ ADC and Programmable Compressed Sensing,” IEEE Journal of Solid- State Circuits, vol. 48, no. 1, pp. 318–328, Jan. 2013
2013
-
[50]
Learned block-based hybrid image compression,
Y . Wu, X. Li, Z. Zhang, X. Jin, and Z. Chen, “Learned block-based hybrid image compression,” ArXiv, vol. abs/2012.09550, 2020
2012 arXiv
-
[51]
Fast Gradient-Based Algorithms for Con- strained Total Variation Image Denoising and Deblurring Problems,
A. Beck and M. Teboulle, “Fast Gradient-Based Algorithms for Con- strained Total Variation Image Denoising and Deblurring Problems,” IEEE Transactions on Image Processing , vol. 18, no. 11, pp. 2419– 2434, Nov. 2009
2009
-
[52]
High-order incremental sigma–delta for compressive sensing and its application to image sen- sors,
W. Guicquero, A. Verdant, and A. Dupret, “High-order incremental sigma–delta for compressive sensing and its application to image sen- sors,” Electronics Letters, vol. 51, no. 19, pp. 1492–1494, 2015
2015
-
[53]
ReconNet: Non-Iterative Reconstruction of Images from Compressively Sensed Measurements,
K. Kulkarni, S. Lohit, P. Turaga, R. Kerviche, and A. Ashok, “ReconNet: Non-Iterative Reconstruction of Images from Compressively Sensed Measurements,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2016
2016
-
[54]
Dr2-net: Deep residual reconstruction network for image compressive sensing,
H. Yao, F. Dai, D. Zhang, Y . Ma, S. Zhang, and Y . Zhang, “Dr2-net: Deep residual reconstruction network for image compressive sensing,” CoRR, vol. abs/1702.05743, 2017. [Online]. Available: http://arxiv.org/abs/1702.05743
2017 arXiv
-
[55]
Variable rate image compres- sion with recurrent neural network,
G. Toderici, S. M. O’Malley, S. J. Hwang, D. Vincent, D. C. Minnen, S. Baluja, M. Covell, and R. Sukthankar, “Variable rate image compres- sion with recurrent neural network,” in ICLR, 2016
2016
-
[56]
Full Resolution Image Compression with Recurrent Neural Networks,
G. Toderici, D. Vincent, N. Johnston, S. J. Hwang, D. Minnen, J. Shor, and M. Covell, “Full Resolution Image Compression with Recurrent Neural Networks,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Jul. 2017
2017
-
[57]
Learning to inpaint for image compression,
M. H. Baig, V . Koltun, and L. Torresani, “Learning to inpaint for image compression,” in NIPS, 2017
2017
-
[58]
Learning convolutional networks for content-weighted image compression,
M. Li, W. Zuo, S. Gu, D. Zhao, and D. Zhang, “Learning convolutional networks for content-weighted image compression,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2018, pp. 3214–3223
2018
-
[59]
Efficient variable rate image compression with multi-scale decomposition network,
C. Cai, L. Chen, X. Zhang, and Z. Gao, “Efficient variable rate image compression with multi-scale decomposition network,” IEEE Transac- tions on Circuits and Systems for Video Technology , vol. 29, no. 12, pp. 3687–3700, 2019
2019
-
[60]
Ensemble learning-based rate-distortion optimization for end-to-end image compression,
Y . Wang, D. Liu, S. Ma, F. Wu, and W. Gao, “Ensemble learning-based rate-distortion optimization for end-to-end image compression,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 31, no. 3, pp. 1193–1207, 2021
2021
-
[61]
Efficient super resolution using binarized neural network,
Y . Ma, H. Xiong, Z. Hu, and L. Ma, “Efficient super resolution using binarized neural network,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , pp. 694–703, 2019
2019
-
[62]
Binarized Neural Network for Single Image Super Resolution,
J. Xin, N. Wang, X. Jiang, J. Li, H. Huang, and X. Gao, “Binarized Neural Network for Single Image Super Resolution,” in Computer Vision – ECCV 2020 . Cham: Springer International Publishing, 2020
2020
-
[63]
Learning Multiple Layers of Features from Tiny Images,
A. Krizhevsky, “Learning Multiple Layers of Features from Tiny Images,” p. 60, 2009. [Online]. Available: https://www.cs.toronto.edu/ ∼kriz/learning-features-2009-TR.pdf
2009
-
[64]
Low bit-width convolutional neural network on rram,
Y . Cai, T. Tang, L. Xia, B. Li, Y . Wang, and H. Yang, “Low bit-width convolutional neural network on rram,”IEEE Transactions on Computer- Aided Design of Integrated Circuits and Systems , vol. 39, no. 7, pp. 1414–1427, 2020
2020
-
[65]
Very Deep Convolutional Networks for Large-Scale Image Recognition,
K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” arXiv:1409.1556 [cs], Apr. 2015
2015 arXiv
-
[66]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–778, 2016
2016
-
[67]
Mobilenets: Efficient convolutional neural networks for mobile vision applications,
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” ArXiv, vol. abs/1704.04861, 2017
2017 arXiv
-
[68]
Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation,
Y . Bengio, N. L ´eonard, and A. Courville, “Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation,” arXiv:1308.3432 [cs], Aug. 2013
2013 arXiv
-
[69]
Deeply-Supervised Nets,
C.-Y . Lee, S. Xie, P. Gallagher, Z. Zhang, and Z. Tu, “Deeply-Supervised Nets,” in Artificial Intelligence and Statistics . Proceedings of Machine Learning Research (PMLR), Feb. 2015
2015
-
[70]
Adam: A Method for Stochastic Optimization,
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” arXiv:1412.6980 [cs], Jan. 2017
2017 arXiv
-
[71]
Adaptive quantization method for cnn with computational-complexity-aware reg- ularization,
K. Nakata, D. Miyashita, J. Deguchi, and R. Fujimoto, “Adaptive quantization method for cnn with computational-complexity-aware reg- ularization,” in 2021 IEEE International Symposium on Circuits and Systems (ISCAS), 2021, pp. 1–5
2021
-
[72]
Differentiable joint pruning and quantization for hardware efficiency,
Y . Wang, Y . Lu, and T. Blankevoort, “Differentiable joint pruning and quantization for hardware efficiency,” in Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XXIX , ser. Lecture Notes in Computer Science, A. Vedald...
2020 doi
-
[73]
Ntire 2017 challenge on single image super-resolution: Dataset and study,
E. Agustsson and R. Timofte, “Ntire 2017 challenge on single image super-resolution: Dataset and study,” in The IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR) Workshops , July 2017
2017
-
[74]
The jpeg 2000 still im- age compression standard,
A. Skodras, C. Christopoulos, and T. Ebrahimi, “The jpeg 2000 still im- age compression standard,” IEEE Signal Processing Magazine , vol. 18, no. 5, pp. 36–58, 2001
2000
-
[75]
Alternating algorithms for total vari- ation image reconstructionfrom random projections,
Y . Xiao, J. Yang, X. Yuan, ,Institute of Applied Mathematics, Henan University, Kaifeng 475004, ,Department of Mathematics, Nanjing Uni- versity, Nanjing, 210093, and ,Department of Mathematics, Hong Kong Baptist University, Hong Kong, “Alternating algorithms for total vari- ...
2012
-
[76]
Context-adaptive entropy model for end- to-end optimized image compression,
J. Lee, S. Cho, and S. Beack, “Context-adaptive entropy model for end- to-end optimized image compression,” in ICLR, 2019
2019
-
[77]
Multi-scale structural similar- ity for image quality assessment,
Z. Wang, E. P. Simoncelli, and A. Bovik, “Multi-scale structural similar- ity for image quality assessment,” in In Signals, Systems and Computers
-
[2004]
Conference Record of the Thirty-Seventh Asilomar Conference , 2003
2003
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.