Pith. sign in

REVIEW 3 major objections 3 minor 66 references

MobiSR: Efficient On-Device Super-Resolution through Heterogeneous Mobile Processors

T0 review · 3 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read MobiSR claims that a difficulty-aware scheduler using total variation can cut on-device super-resolution latency by 2.13x to 4.79x while staying within a user-specified PSNR drop.

desk verdict MobiSR is a genuinely useful mobile SR scheduling system, but the abstract's 2.13x/4.79x speedups are unconstrained averages; the equal-quality speedups in Fig. 10 are at most ~1.94x. read the letter →

arxiv 1908.07985 v1 pith:DGU2ZZOT submitted 2019-08-21 cs.CV cs.DC

classification cs.CVcs.DC
keywords super-resolutionon-deviceinferenceheterogeneousmobileprocessorsdifficulty-awareschedulingtotalvariationmodelcompressionPSNR-latencytrade-offdesignspaceexploration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

MobiSR sets out to show that on-device super-resolution can be made much faster by treating image patches differently: easy patches go to a large, accurate model, while hard patches go to a compact, low-precision model. The paper's central assertion is that a patch's total variation is a usable run-time proxy for its upscaling difficulty, so a simple threshold can dispatch patches without exceeding a user-set average PSNR budget. The framework automatically searches compressed variants of a supplied super-resolution network, prunes to Pareto-optimal models per compute engine, and tunes the threshold to minimize latency subject to the error tolerance. On a mobile SoC with CPU, GPU, and DSP, the resulting two-model designs achieve an average speedup of 2.13x over parallel difficulty-unaware mappings and 4.79x over single compute engine implementations. If right, this gives mobile zoom and image-upscaling applications a practical path to local processing with bounded quality loss.

What carries the argument

The load-bearing mechanism is the total-variation threshold scheduler. Total variation (TV) is the sum of absolute intensity differences between neighboring pixels in a patch, and the scheduler computes it per patch at run time; if $TV(p) \le TV_{\mathrm{thr}}$ the patch is sent to the larger model on the PSNR-preserving engines (CPU/GPU), otherwise to the compact model on the low-precision engine (DSP), with load balancing across engines and an overflow rule that lets hard patches fall back to the large model when the DSP is busy. The threshold $TV_{\mathrm{thr}}$ is not hand-set: a calibration set is used to estimate the domain's TV range, and the optimization over model pairs and thresholds is driven by an analytical performance model that estimates per-engine latency from measured per-patch execution times. The design space itself is generated by model transformations—residual bottlenecks, group, depthwise-separable, and separable convolutions, inverted residuals, channel shuffle, and channel split—which are pruned per engine to the Pareto front.

What would settle it

Run both candidate models on every patch of a held-out dataset that differs from the calibration set, compute per-patch PSNR against the full-precision large model, and count high-TV patches where the compact model's PSNR drop exceeds the user tolerance; if this fraction is large, the TV-threshold scheduler cannot meet its quality constraint, and the reported speedup at fixed quality would not transfer.

Watch

Extended reading notes

Core claim

The discovery is that image patches differ systematically in how much a large super-resolution model outperforms a compact one, and that this difference tracks total variation. For low-TV (easy) patches the large model wins by a meaningful margin, so those patches should be sent to it; for high-TV (hard) patches the two models are nearly equally weak, so sending them to a fast compact model costs little quality and saves time. MobiSR turns this into a deployable system: it generates a family of compressed models from a user-supplied network, measures their latency on each compute engine, keeps only Pareto-optimal (model, engine) pairs, and then selects two models and a TV threshold that minimize estimated latency while keeping average PSNR within a user-specified drop. The quantitative claim is an average speedup of 2.13x over highly optimized parallel difficulty-unaware mappings and 4.79x over highly optimized single-engine implementations, measured on one mobile SoC.

Load-bearing premise

Total variation of a patch is a reliable run-time proxy for how hard that patch is to upscale, and on hard patches the compact model's quality loss is small enough that the user's average PSNR budget is still met.

Editorial extensions

If this is right

  • For a fixed error tolerance, an on-device SR system can run faster by splitting the image by difficulty rather than running one model on all patches.
  • The framework's reported speedups are 2.13x over parallel difficulty-unaware mappings and 4.79x over single-engine implementations, with larger speedups as the allowed PSNR drop increases.
  • Because the optimization is parametrized by the target SoC's compute engines, the same procedure applies to newer chips with neural accelerators, not only to CPU/GPU/DSP designs.
  • The Pareto-front pruning and the constraint that the fast model be more compact than the accurate model keep the design search small enough to be exhaustive.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the scheduler's quality guarantee is average PSNR, not per-patch or worst-case quality; applications sensitive to artifacts in individual patches would need a stricter dispatch rule.
  • Beyond the paper: the measured speedups come from one mobile SoC; on newer heterogeneous chips the relative latencies of CPU, GPU, and DSP change, so the optimal model pair and threshold would shift, though the search and scheduling procedure should transfer.
  • Beyond the paper: a natural testable extension is to make the threshold content-adaptive per image or to learn it from a small labeled calibration set instead of a single global value, which could improve the quality-latency frontier.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper presents MobiSR, a framework for on-device super-resolution that combines two compressed SR models with a runtime scheduler. The scheduler uses the total variation (TV) of each image patch as a proxy for upscaling difficulty, dispatching easy patches to a larger, more accurate model on the CPU/GPU and hard patches to a compact, quantized model on the DSP, with load balancing. Offline, the framework explores model compression transformations and selects a model pair and a TV threshold to minimize latency subject to a user-specified PSNR drop constraint, formalized as the optimization problem in Eq. (12). The authors report average speedups of 2.13x over optimized parallel difficulty-unaware mappings and 4.79x over single-engine implementations on a Snapdragon 845 board, with additional measurements across Set5, Set14, B100, and Urban100.

Significance. The paper addresses a practically relevant systems problem: efficient on-device super-resolution under a quality constraint. Its strengths are concrete: the evaluation uses real measurements on the Snapdragon 845 via SNPE, the design-space exploration with compression transformations is clearly described, and Fig. 10 honestly reports speedups as a function of allowed PSNR degradation. If the central speedup claim is properly connected to the quality constraint, the work would be a useful contribution to mobile systems and embedded deep learning. However, the headline numbers in the abstract and Table 3 are not computed under the equal-quality protocol used in Fig. 10, and the TV-based difficulty proxy is validated only at the image level while being applied at the patch level. These issues affect the central claim and require revision.

major comments (3)
  1. [Abstract; Table 3; Section 4.4, Fig. 10] The headline speedups of 2.13x and 4.79x are not supported by the equal-quality evaluation. Table 3 reports speedup ranges and averages over TV threshold values without reporting the PSNR drop at each operating point, so the averages mix configurations that violate the Eq. (12) constraint with those that satisfy it. In contrast, Fig. 10, which compares systems at the same PSNR drop, shows maximum speedups of 47%, 78%, 94%, and 29% on Set5, Set14, B100, and Urban100, respectively. The abstract should either report speedups under the PSNR constraint or clearly label the Table 3 numbers as unconstrained operating points.
  2. [Section 3.3, Eq. (9), Fig. 4] The TV-based difficulty criterion is validated at the image level on DIV2K, but the scheduler operates on individual patches. The paper does not provide patch-level evidence that high-TV patches are indeed nearly equally hard for the large and compact models, nor does it describe a held-out calibration protocol for selecting TV_thr. Because the scheduler's quality guarantee depends on this transfer, the authors should add a patch-level analysis (e.g., PSNR difference vs. TV for patches) and an explicit calibration/validation split for threshold selection.
  3. [Section 3.4, Eq. (11); Algorithm 1] The performance model in Eq. (11) assumes that patches are assigned to m1 or m2 solely by the TV threshold, but Algorithm 1 allows hard patches to be processed by m1 when the DSP is oversubscribed. This means the analytic latency estimate can understate the running time by assuming more DSP utilization than the scheduler actually provides. The authors should either incorporate the load-balancing policy into the performance model or justify that the discrepancy is negligible for the evaluated configurations.
minor comments (3)
  1. [Fig. 9 caption] The caption reads 'SDM854' but the platform is Snapdragon 845 (SDM845); this typo should be corrected.
  2. [Section 4.4, Fig. 10] The x-axis of Fig. 10 is described only as 'error degradation'; the exact PSNR drop values used for each interval are not stated in the text, making it hard to assess which operating points are of practical interest.
  3. [Table 3] The table would be easier to interpret if it also reported the PSNR drop of each configuration at the listed speedups, or if the speedups were restricted to configurations satisfying the Eq. (12) constraint.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; MobiSR is an empirically benchmarked systems paper with an independently measured central claim.

full rationale

MobiSR is an empirical systems paper, not a derivation: the central speedup claim is measured on-device against external SR datasets and baselines, and the TV-based difficulty proxy is validated with DIV2K scatter plots rather than assumed from the target result. The optimizer in Eq. (12) tunes the TV threshold to meet a user-specified PSNR error tolerance, and Fig. 10 compares MobiSR against single-model baselines under matched PSNR-drop budgets, so the equal-quality comparison is not definitionally tied to the fitted threshold. The reader and skeptic concerns about Table 3 averaging speedups over thresholds without enforcing the PSNR constraint are validity or reporting concerns, not circularity: no equation or fitted parameter is renamed as a prediction. Self-citations such as [2] and [52] are background references and are not load-bearing for the paper's main claim. No load-bearing step reduces by construction to its own input, and no uniqueness or ansatz is imported via self-citation.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new physical or mathematical entities. The free parameters are the TV threshold, compression ratios, and patch size; the central speedup claim depends on these choices. The load-bearing assumptions are empirical correlations between total variation and SR quality, which are plausible but only shown at image level on DIV2K.

free parameters (3)
  • TV_thr (total-variation threshold) = Per-dataset values not reported numerically; shown as sweeps in Fig. 9 and Table 3
    The scheduler's hard/easy cutoff is selected to meet the PSNR budget; it is a fitted design parameter of the optimization in Eq. (12).
  • Compression hyperparameters (r, g, e) = r=2, r=4, g=4, g=16, e=2
    The candidate model space is generated by hand-chosen ratios from the compression literature (Table 4), so the final selected pair and the reported speedups depend on this manually chosen candidate set.
  • Patch size = 90x160
    Patch dimensions are fixed by hand in Section 4.1 and affect both TV calculations and load-balancing granularity.
assumptions (4)
  • domain assumption Total variation is a reliable proxy for upscaling difficulty.
    Section 3.3, Eq. (9) and Figs. 3 and 5: used to classify patches as easy or hard; the supporting plot is image-level, on DIV2K.
  • domain assumption For hard-to-upscale patches, the large and compact models have similar PSNR, so the compact model can serve hard patches without breaking the average PSNR budget.
    Section 3.3, Fig. 4: this empirical observation justifies sending hard patches to the weaker model; it is dataset-specific and not independently validated.
  • domain assumption Image-level TV-PSNR behavior transfers to individual patches at run time.
    Algorithm 1 applies the TV criterion per patch, but the paper validates the relationship only at image level on DIV2K.
  • domain assumption SNPE latency measurements on the Open-Q 845 board are representative of the target mobile platform class.
    Section 4.1: all latency numbers come from one SDK version and one development board; generalization to other SoCs is asserted, not measured.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MobiSR: Efficient On-Device Super-Resolution through Heterogeneous Mobile Processors." pith.science (2026). https://pith.science/paper/DGU2ZZOT

@misc{pith2026190807985,
  author       = {Pith},
  title        = {Pith review of: MobiSR: Efficient On-Device Super-Resolution through Heterogeneous Mobile Processors},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DGU2ZZOT}},
  note         = {Machine review of arXiv:1908.07985}
}
read the original abstract

In recent years, convolutional networks have demonstrated unprecedented performance in the image restoration task of super-resolution (SR). SR entails the upscaling of a single low-resolution image in order to meet application-specific image quality demands and plays a key role in mobile devices. To comply with privacy regulations and reduce the overhead of cloud computing, executing SR models locally on-device constitutes a key alternative approach. Nevertheless, the excessive compute and memory requirements of SR workloads pose a challenge in mapping SR networks on resource-constrained mobile platforms. This work presents MobiSR, a novel framework for performing efficient super-resolution on-device. Given a target mobile platform, the proposed framework considers popular model compression techniques and traverses the design space to reach the highest performing trade-off between image quality and processing speed. At run time, a novel scheduler dispatches incoming image patches to the appropriate model-engine pair based on the patch's estimated upscaling difficulty in order to meet the required image quality with minimum processing latency. Quantitative evaluation shows that the proposed framework yields on-device SR designs that achieve an average speedup of 2.13x over highly-optimized parallel difficulty-unaware mappings and 4.79x over highly-optimized single compute engine implementations.

Figures

Figures reproduced from arXiv: 1908.07985 by the authors.

Figure 2
Figure 2. MobiSR’s processing flow. optimizations, but also on the model and the level of quanti￾zation involved. Therefore, a key challenge in accelerating SR models is selecting appropriate convolution-approximation techniques based on both their impact on the accuracy of the given model and their efficient mapping on the available compute engines. 3 MobiSR In this section, we present the high-level flow of MobiSR followed … view at source ↗
Figure 2
Figure 2. The framework is supplied with a high-level descrip [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Low-resolution images along with their TV [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (5 more)
Figure 5
Figure 5. Figure 5: TV of low-resolution images from the DIV2K [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: We reduce the number of feature maps in RCAN to improve processing speed. Conv(S,D,K) is a K×K convolution with S input and D output feature maps. Implementation Details. All SR designs presented in this section were run on SDM845 using the high-performance profile of …
Figure 7
Figure 7. Figure 7: PSNR vs CPU latency of MobiSR-generated models on SDM845 (×4 upscaling on Urban100). model,mref, as RCAN yields the state-of-the-art performance based on PSNR/SSIM among large-scale SR models. In order for RCAN to be comparable to existing state-of-the-art mo￾bile SR m…
Figure 9
Figure 9. Figure 9: Achieved PSNR and measured performance as a function of TV on SDM854. [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 10
Figure 10. Figure 10: MobiSR’s speedup as a function of error degradation. [PITH_FULL_IMAGE:figures/full_fig_p012_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

66 extracted references · 64 canonical work pages

  1. [1]

    Namhyuk Ahn, Byungkon Kang, and Kyung-Ah Sohn. 2018. Fast, Accurate, and Lightweight Super-Resolution with Cascading Residual Network. In ECCV

  2. [2]

    Ve- nieris, and Nicholas D

    Mario Almeida, Stefanos Laskaridis, Ilias Leontiadis, Stylianos I. Ve- nieris, and Nicholas D. Lane. 2019. EmBench: Quantifying Performance Variations of Deep Neural Networks Across Modern Commodity De- vices. In The 3rd International Workshop on Deep Learning for Mobile Systems and Applications (EMDL ’19) . 1–6

  3. [3]

    Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural Machine Translation by Jointly Learning to Align and Translate. In International Conference on Learning Representations (ICLR)

  4. [4]

    Marco Bevilacqua, Aline Roumy, Christine Guillemot, and Marie line Alberi Morel. 2012. Low-Complexity Single-Image Super-Resolution based on Nonnegative Neighbor Embedding. In British Machine Vision Conference (BMVC)

  5. [5]

    A. M. Caulfield, E. S. Chung, A. Putnam, H. Angepat, J. Fowers, M. Haselman, S. Heil, M. Humphrey, P. Kaur, J. Y. Kim, D. Lo, T. Massengill, K. Ovtcharov, M. Papamichael, L. Woods, S. Lanka, D. Chiou, and D. Burger. 2016. A Cloud-Scale Acceleration Architecture. In 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). 1–13

  6. [6]

    Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. 2016. Image Super-Resolution Using Deep Convolutional Networks. IEEE Transactions on Pattern Analysis and Machine Intelligence 38 (2016), 295–307

  7. [7]

    Chao Dong, Chen Change Loy, and Xiaoou Tang. 2016. Accelerating the Super-Resolution Convolutional Neural Network. In ECCV

  8. [8]

    Fowers, K

    J. Fowers, K. Ovtcharov, M. Papamichael, T. Massengill, M. Liu, D. Lo, S. Alkalay, M. Haselman, L. Adams, M. Ghandi, S. Heil, P. Patel, A. Sapek, G. Weisz, L. Woods, S. Lanka, S. K. Reinhardt, A. M. Caulfield, E. S. Chung, and D. Burger. 2018. A Configurable Cloud-Scale DNN Pro- cessor for Real-Time AI. In 2018 ACM/IEEE 45th Annual International Symposium...

Show all 66 references
  1. [9]

    Ido Freeman, Lutz Roese-Koerner, and Anton Kummert. 2018. Effnet: An Efficient Structure for Convolutional Neural Networks. In IEEE International Conference on Image Processing (ICIP) . 6–10

  2. [10]

    Ian Goodfellow et al. 2014. Generative Adversarial Nets. In Advances in Neural Processing Systems

  3. [11]

    Song Han, Huizi Mao, and William J Dally. 2016. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quanti- zation and Huffman Coding. In International Conference on Learning Representations (ICLR)

  4. [12]

    Song Han, Jeff Pool, John Tran, and William Dally. 2015. Learning both Weights and Connections for Efficient Neural Network. In Advances in Neural Information Processing Systems 28 . 1135–1143

  5. [13]

    Seungyeop Han, Haichen Shen, Matthai Philipose, Sharad Agar- wal, Alec Wolman, and Arvind Krishnamurthy. 2016. MCDNN: An Approximation-Based Execution Framework for Deep Stream Process- ing Under Resource Constraints. In Proceedings of the 14th Annual International Conference ...

  6. [14]

    Kim Hazelwood, Sarah Bird, David Brooks, Soumith Chintala, Utku Diril, Dmytro Dzhulgakov, Mohamed Fawzy, Bill Jia, Yangqing Jia, Aditya Kalro, James Law, Kevin Lee, Jason Lu, Pieter Noordhuis, Misha Smelyanskiy, Liang Xiong, and Xiaodong Wang. 2018. Applied Ma- chine Learning ...

  7. [15]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR) . 770–778

  8. [16]

    Z. He, H. Huang, M. Jiang, Y. Bai, and G. Luo. 2018. FPGA-Based Real- Time Super-Resolution System for Ultra High Definition Videos. In 2018 IEEE 26th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM) . 181–188

  9. [17]

    Gopalakrishna Hegde, Siddhartha, and Nachiket Kapre. 2017. Caf- fePresso: Accelerating Convolutional Networks on Embedded SoCs. ACM Trans. Embed. Comput. Syst.17, 1, Article 15 (Nov. 2017), 26 pages

  10. [18]

    Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean. 2015. Distilling the Knowledge in a Neural Network. In NIPS Deep Learning and Represen- tation Learning Workshop

  11. [19]

    Howard et al

    Andrew G. Howard et al. 2017. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications.CoRR abs/1704.04861 (2017)

  12. [20]

    Gibbons, and Onur Mutlu

    Kevin Hsieh, Ganesh Ananthanarayanan, Peter Bodik, Shivaram Venkataraman, Paramvir Bahl, Matthai Philipose, Phillip B. Gibbons, and Onur Mutlu. 2018. Focus: Querying Large Video Datasets with Low Latency and Low Cost. In Proceedings of the 12th USENIX Conference on Operating S...

  13. [21]

    Gao Huang, Danlu Chen, Tianhong Li, Felix Wu, Laurens van der Maaten, and Kilian Weinberger. 2018. Multi-Scale Dense Networks for Resource Efficient Image Classification. In International Conference on Learning Representations (ICLR)

  14. [22]

    Huang, A

    J. Huang, A. Singh, and N. Ahuja. 2015. Single image super-resolution from transformed self-exemplars. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

  15. [23]

    Zheng Hui, Xiumei Wang, and Xinbo Gao. 2018. Fast and Accurate Single Image Super-Resolution via Information Distillation Network. In CVPR. 723–731

  16. [24]

    Loc Huynh, Rajesh Krishna Balan, and Youngki Lee. 2016. DeepSense: A GPU-based Deep Convolutional Neural Network Framework on Commodity Mobile Devices. In WearSys@MobiSys

  17. [25]

    Andrey Ignatov et al . 2018. PIRM Challenge on Perceptual Image Enhancement on Smartphones: Report. In ECCV Workshops

  18. [26]

    Andrey Ignatov, Radu Timofte, William Chou, Ke Wang, Max Wu, Tim Hartley, and Luc Van Gool. 2018. AI Benchmark: Running Deep Neural Networks on Android Smartphones. In The European Conference on Computer Vision (ECCV) Workshops

  19. [27]

    Justin Johnson, Alexandre Alahi, and Li Fei-Fei. 2016. Perceptual Losses for Real-Time Style Transfer and Super-Resolution. In ECCV

  20. [28]

    Daniel Kang, John Emmons, Firas Abuzaid, Peter Bailis, and Matei Zaharia. 2017. NoScope: Optimizing Neural Network Queries over Video at Scale. Proc. VLDB Endow. 10, 11 (Aug. 2017), 1586–1597

  21. [29]

    Yigitcan Kaya, Sanghyun Hong, and Tudor Dumitras. [n. d.]. Shallow- Deep Networks: Understanding and Mitigating Network Overthinking. In International Conference on Machine Learning (ICML)

  22. [30]

    Jiwon Kim, Jung Kwon Lee, and Kyoung Mu Lee. 2016. Accurate Image Super-Resolution Using Very Deep Convolutional Networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 1646–1654

  23. [31]

    Y. Kim, J. Choi, and M. Kim. 2018. A Real-Time Convolutional Neural Network for Super-Resolution on FPGA with Applications to 4K UHD 60 fps Video Services. IEEE Transactions on Circuits and Systems for Video Technology (2018)

  24. [32]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Sto- chastic Optimization. In International Conference on Learning Repre- sentations (ICLR)

  25. [33]

    Kouris, S

    A. Kouris, S. I. Venieris, and C. Bouganis. 2018. CascadeCNN: Pushing the Performance Limits of Quantisation in Convolutional Neural Net- works. In 2018 28th International Conference on Field Programmable Logic and Applications (FPL). 155–1557. https://doi.org/10.1109/FPL. 2018.00034

  26. [34]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. Ima- geNet Classification with Deep Convolutional Neural Networks. In NIPS

  27. [35]

    N. D. Lane, S. Bhattacharya, P. Georgiev, C. Forlivesi, L. Jiao, L. Qendro, and F. Kawsar. 2016. DeepX: A Software Accelerator for Low-Power Deep Learning Inference on Mobile Devices. In 2016 15th ACM/IEEE International Conference on Information Processing in Sensor Networks (...

  28. [36]

    N. D. Lane, S. Bhattacharya, A. Mathur, P. Georgiev, C. Forlivesi, and F. Kawsar. 2017. Squeezing Deep Learning into Mobile and Embedded Devices. IEEE Pervasive Computing 16, 3 (2017), 82–88

  29. [37]

    Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, and Kyoung Mu Lee. 2017. Enhanced Deep Residual Networks for Single Image Super- Resolution. In IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR) Workshops

  30. [38]

    Chao Ma, Chih-Yuan Yang, Xiaokang Yang, and Ming-Hsuan Yang

  31. [39]

    Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun. 2018. ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design. In The European Conference on Computer Vision (ECCV)

  32. [40]

    Martin, C

    D. Martin, C. Fowlkes, D. Tal, and J. Malik. 2001. A database of human segmented natural images and its application to evaluating segmenta- tion algorithms and measuring ecological statistics. In IEEE Interna- tional Conference on Computer Vision (ICCV)

  33. [41]

    Completely Blind

    Anish Mittal, Rajiv Soundararajan, and Alan C. Bovik. 2013. Making a “Completely Blind" Image Quality Analyzer. IEEE Signal Processing Letters 20 (2013), 209–212

  34. [42]

    Seyyed Salar Latifi Oskouei, Hossein Golestani, Matin Hashemi, and Soheil Ghiasi. 2016. CNNdroid: GPU-Accelerated Execution of Trained Deep Convolutional Neural Networks on Android. In ACM Multime- dia

  35. [43]

    Jiantao Qiu, Jie Wang, Song Yao, Kaiyuan Guo, Boxun Li, Erjin Zhou, Jincheng Yu, Tianqi Tang, Ningyi Xu, Sen Song, Yu Wang, and Huazhong Yang. 2016. Going Deeper with Embedded FPGA Plat- form for Convolutional Neural Network. In Proceedings of the 2016 ACM/SIGDA International ...

  36. [44]

    Rudin, Stanley Osher, and Emad Fatemi

    Leonid I. Rudin, Stanley Osher, and Emad Fatemi. 1992. Nonlinear Total Variation Based Noise Removal Algorithms. Phys. D 60, 1-4 (Nov. 1992), 259–268

  37. [45]

    Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. 2018. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

  38. [46]

    Haichen Shen, Seungyeop Han, Matthai Philipose, and Arvind Krish- namurthy. 2017. Fast Video Classification via Adaptive Cascading of Deep Models. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

  39. [47]

    Laurent Sifre and Stéphane Mallat. 2014. Rigid-motion scattering for image classification. PhD thesis, Ph. D. thesis 1 (2014), 3

  40. [48]

    Mingcong Song, Kan Zhong, Jiaqi Zhang, Yang Hu, Duo Liu, Weigong Zhang, Jing Wang, and Tao E Li. 2018. In-Situ AI: Towards Autonomous and Incremental Deep Learning for IoT Systems. 2018 IEEE Interna- tional Symposium on High Performance Computer Architecture (HPCA) (2018), 92–103

  41. [49]

    Teerapittayanon, B

    S. Teerapittayanon, B. McDanel, and H. T. Kung. 2016. BranchyNet: Fast inference via early exiting from deep neural networks. In 2016 23rd International Conference on Pattern Recognition (ICPR). 2464–2469

  42. [50]

    Timofte et al

    R. Timofte et al. 2017. NTIRE 2017 Challenge on Single Image Super- Resolution: Methods and Results. In IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)

  43. [51]

    Venieris and Christos-Savvas Bouganis

    Stylianos I. Venieris and Christos-Savvas Bouganis. 2018. fpgaConvNet: Mapping Regular and Irregular Convolutional Neural Networks on FPGAs. IEEE Transactions on Neural Networks and Learning Systems 30 (2018), 326–342

  44. [52]

    Stylianos I Venieris, Alexandros Kouris, and Christos-Savvas Bouganis

  45. [53]

    Thang Vu, Cao Van Nguyen, Trung Xuan Pham, Tung Minh Luu, and Chang Dong Yoo. 2018. Fast and Efficient Image Quality Enhance- ment via Desubpixel Convolutional Neural Networks. In The European Conference on Computer Vision (ECCV) Workshops

  46. [54]

    Bovik, Hamid R

    Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli

  47. [55]

    Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He

    Saining Xie, Ross B. Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. 2017. Aggregated Residual Transformations for Deep Neural Net- works. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

  48. [56]

    Huang, and Yi Ma

    Jianchao Yang, John Wright, Thomas S. Huang, and Yi Ma. 2010. Image Super-resolution via Sparse Representation. Trans. Img. Proc. 19, 11 (2010), 2861–2873

  49. [57]

    Tien-Ju Yang, Yu-Hsin Chen, and Vivienne Sze. 2017. Designing Energy-Efficient Convolutional Neural Networks Using Energy-Aware Pruning. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

  50. [58]

    Dong-Qing Zhang. 2018. clcNet: Improving the Efficiency of Convo- lutional Neural Network Using Channel Local Convolutions. IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2018)

  51. [59]

    Freedman

    Haoyu Zhang, Ganesh Ananthanarayanan, Peter Bodik, Matthai Phili- pose, Paramvir Bahl, and Michael J. Freedman. 2017. Live Video Ana- lytics at Scale with Approximation and Delay-tolerance. InProceedings of the 14th USENIX Conference on Networked Systems Design and Im- plement...

  52. [60]

    Kai Zhang, Wangmeng Zuo, and Lei Zhang. 2018. Learning a Single Convolutional Super-Resolution Network for Multiple Degradations. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

  53. [61]

    Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. 2018. Shuf- fleNet: An Extremely Efficient Convolutional Neural Network for Mo- bile Devices. In IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR)

  54. [62]

    Yulun Zhang, Kunpeng Li, Kai Li, Lichen Wang, Bineng Zhong, and Yun Fu. 2018. Image Super-Resolution Using Very Deep Residual Channel Attention Networks. In European Conference on Computer Vision (ECCV)

  55. [63]

    Yulun Zhang, Yapeng Tian, Yu Kong, Bineng Zhong, and Yun Fu. 2018. Residual Dense Network for Image Super-Resolution. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

  56. [2004]

    IEEE Transactions on Image Processing 13 (2004), 600–612

    Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing 13 (2004), 600–612

  57. [2017]

    Computer Vision and Image Understanding 158 (2017), 1–16

    Learning a No-Reference Quality Metric for Single-Image Super- Resolution. Computer Vision and Image Understanding 158 (2017), 1–16

  58. [2018]

    In 2nd International Workshop on Embedded and Mobile Deep Learning (EMDL)

    Deploying Deep Neural Networks in the Embedded Space. In 2nd International Workshop on Embedded and Mobile Deep Learning (EMDL)

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.