Pith. sign in

REVIEW 1 major objections 6 minor 153 references

A Survey on Efficiency Optimization Techniques for DNN-based Video Analytics: Process Systems, Algorithms, and Applications

T0 review · 1 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This survey proposes a four-layer, bottom-up taxonomy of efficiency optimization techniques for DNN-based video analytics, covering storage, computing systems, DNN algorithms, and applications, and positions itself as the first systematic…

desk verdict A useful four-layer survey whose central technique-to-paper mapping is undercut by a confirmed mis-citation in Table 2; fix the audit and it earns its place as a reference. read the letter →

arxiv 2507.15628 v1 pith:P2I5UG6B submitted 2025-07-21 cs.CV

classification cs.CV
keywords videoanalyticsefficiencyoptimizationdeepneuralnetworksedgecomputingcloudmodelcompressionsurveillancesurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that existing reviews of DNN-based video analytics emphasize accuracy or cover only one stage of the pipeline, leaving the efficiency dimension unsystematized. It organizes the field into four stacked layers — storage support, computing system, DNN algorithm, and application — and reviews techniques within each from the bottom up. The intended payoff is a map that lets researchers and engineers locate the optimization approach they need and see open problems. The survey claims this bottom-up, cross-layer organization is its main contribution.

What carries the argument

The organizing device is a four-layer bottom-up taxonomy of efficiency optimization techniques, defined as the stacked classification of storage support, computing system, DNN algorithm, and application layers. It carries the argument by turning a scattered literature into a structured map: each technique is assigned to a layer, with temporal and spatial optimization presented as orthogonal axes at the algorithm layer. The taxonomy also drives the survey's forward-looking discussion of challenges and open issues.

What would settle it

Two checks settle the central claims: a literature search for an earlier survey that already organizes efficiency techniques across storage, computing, algorithm, and application layers bottom-up; and a full audit of the citation-to-technique mappings in all tables, looking for rows where the cited papers do not contain the described technique (as the Video Frame Filtering row does for FilterForward, citing [34,36,37] while the text cites [63]).

Watch

Extended reading notes

Core claim

The central claim is that efficiency optimization for DNN-based video analytics can be systematically understood through a four-layer framework: the storage supporting layer (hybrid memory and video database systems), the computing system layer (edge, edge-cloud collaborative, and cloud cluster techniques such as model compression, frame filtering, encoding compression, task partitioning, and scheduling), the DNN algorithm layer (spatial techniques at pixel, region, and resolution granularity and temporal techniques at feature, frame, and sequence granularity), and the application layer (transportation, retail, and industrial security). The paper presents this as the first bottom-up survey devoted to efficiency rather than accuracy, and it identifies open challenges in AI hardware support, video foundation models, 6G-enabled edge-cloud scheduling, and new application scenarios.

Load-bearing premise

The survey's value depends on its citations being correct: if the table-to-paper mappings are wrong, readers cannot trust the taxonomy as a reference.

Editorial extensions

If this is right

  • A reader can classify any efficiency technique for DNN video analytics into one of the four layers and see adjacent techniques it can combine with.
  • Temporal and spatial algorithm optimizations are claimed to be orthogonal, so a system can stack frame sampling with region-level attention for compounded gains.
  • The efficiency-focused, bottom-up organization fills a gap left by accuracy-oriented and stage-specific surveys, giving future work a common reference structure.
  • The open challenges named — hardware-software co-evolution, video foundation model training cost, 6G-enabled scheduling, and new application domains — define a near-term research agenda.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The four-layer taxonomy could be reused as a classification scheme for new papers, giving the field a shared vocabulary for efficiency work.
  • The paper's own Table 2 misattributes FilterForward to quantization papers [34,36,37] in one row while the text cites [63]; if similar errors appear elsewhere, the survey's reliability as a reference map weakens.
  • Cross-layer interactions (for example, how storage layout affects algorithm-level frame sampling) are underexplored; a testable extension would be a system that co-optimizes two layers rather than one.
  • The layered stack suggests a coverage metric: one could measure how many layers a deployed system optimizes and test whether efficiency gains add up as lower-layer support is added.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 6 minor

Summary. This manuscript presents a survey of efficiency optimization techniques for DNN-based video analytics. It organizes the area into four layers—storage supporting, computing system, DNN algorithm, and application—and claims that this bottom-up, efficiency-focused organization distinguishes it from prior accuracy-oriented or stage-specific surveys. The survey reviews storage media and video storage systems, edge and edge-cloud computing techniques (frame filtering, video encoding, task partition, scheduling), model compression and configuration optimization, spatial and temporal algorithm-level optimizations, and applications in transportation, retail, and industrial security. It closes with a discussion of open challenges covering AI hardware, video foundation models, 6G-enabled scheduling, and new application scenarios.

Significance. If its reference mappings and tables are reliable, this survey would be a useful and reasonably comprehensive reference. The four-layer bottom-up organization is a sensible structuring device that goes beyond stage-specific surveys, and the breadth of coverage—about 146 references spanning systems, algorithms, and applications—is a genuine strength. The challenges section is forward-looking and mentions concrete directions rather than generic advice. However, the survey's practical value rests on the accuracy of its citation-to-technique mapping, and at least one confirmed error in Table 2 undermines that value. After a full audit and correction of the tables and reference attributions, the survey could serve as a standard entry point to efficiency optimization in DNN-based video analytics.

major comments (1)
  1. [Table 2, 'Video Frame Filtering' row; cf. §3.2.1] The row for FilterForward cites references [34,36,37], but the running text in §3.2.1 and the bibliography identify FilterForward as reference [63], while [34], [36], and [37] are parameter-quantization papers. A reader using Table 2 to locate the source of edge-side frame filtering would be sent to three unrelated quantization papers. Since Table 2 is one of the paper's primary compressed reference artifacts, this is a load-bearing error for the survey's reliability as a map from techniques to sources. I request a full audit of every row of every table against the text and reference list, not just a one-cell correction.
minor comments (6)
  1. [Running heads (pp. 111:2–111:25)] The running head on every page reads 'Trovato et al.', which does not match the author list; this appears to be a template artifact and should be corrected.
  2. [Fig. 1, Fig. 2, §3.1.1] The figures contain the typos 'Collaberative' and 'Optmization', and §3.1.1 contains 'effectivly'; these should be fixed in a careful copyedit.
  3. [§3.1.4 and §5.3.2] The device name is written as 'NSC2' in §3.1.4 but as 'NCS2' in §5.3.2; the spelling should be made consistent.
  4. [§3.1.1, Network Pruning] Reference [43] is used both to describe 'adaptive channel pruning technique based on naive Bayesian inference' and, later, to describe pruning coupled with Tucker decomposition; the bibliographic entry for [43] does not make the Bayesian component evident, so the authors should clarify the connection or correct the citation.
  5. [§1 and §6] The claimed distinction from prior surveys, especially [23] which also targets deep-learning-driven edge video analytics, would be more convincing with an explicit comparison of coverage; a short paragraph or table stating what each prior survey covers and what this survey adds would make the novelty claim checkable.
  6. [References [59] and [62]] These web-resource references lack retrieval dates and, in one case, a clear publication year; journal style normally requires access dates for online resources.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: this is a descriptive literature survey with no derivations, fitted parameters, or predictions; the only author self-citation (DVQShare [109]) is descriptive and non-load-bearing, and the Table 2 attribution error for FilterForward is a correctness/reliability defect, not a circular step.

full rationale

The paper is a survey with no formal derivation chain: it contains no equations, no fitted parameters, and no empirical predictions. Its central claim is organizational — that it summarizes efficiency optimization techniques for DNN-based video analytics bottom-up across four layers (storage supporting, computing system, DNN algorithm, application). A survey taxonomy is a categorization choice, not a result derived from its inputs, so there is nothing for the claim to reduce to by construction. The only author self-citation is DVQShare [109] (Fu, Tang, Yu, Li — all present authors) in Section 5.1.1, described as a system for concurrent DNN-based video queries. This citation is descriptive, appears once among roughly 130 references, and no section, table, or conclusion depends on DVQShare's specific properties; it is therefore a minor self-citation that is not load-bearing. The manuscript does contain a genuine reliability defect flagged by the reviewer: Table 2's 'Video Frame Filtering' row cites '[34,36,37]' for FilterForward, while Section 3.2.1 and the bibliography identify FilterForward as [63], and [34], [36], [37] are quantization works, not frame filtering systems. This mis-attribution undermines the survey's role as a trustworthy technique-to-paper map and suggests the tables were not fully checked against the primary literature, but it is a factual attribution error rather than a circular step — no claim in the paper is made equivalent to its own input by this error. The running header 'Trovato et al.' is a template artifact with no bearing on circularity. Overall, the survey's substantive content (the descriptions and the four-layer organization) is independent of its own prior work and of any fitted quantity, so the appropriate circularity finding is 'no significant circularity' at the low end of the 0-2 band.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The survey introduces no free parameters, no formal axioms, and no invented entities. Its only assumptions concern the accuracy of the literature summaries and the suitability of the organizing taxonomy.

assumptions (2)
  • domain assumption The cited papers are accurately represented in the survey summaries and tables.
    The survey's value rests on faithful reporting of each referenced system. Table 2 contains a clear mis-citation (FilterForward assigned references [34,36,37] instead of [63]), so this assumption is partially violated.
  • domain assumption The four-layer bottom-up taxonomy (storage, computing system, DNN algorithm, application) is a natural and non-overlapping organization of the field.
    The paper asserts in Section 1 and Fig.1 that these layers comprehensively cover efficiency optimization; the adequacy of this partitioning is not proven.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey on Efficiency Optimization Techniques for DNN-based Video Analytics: Process Systems, Algorithms, and Applications." pith.science (2026). https://pith.science/paper/P2I5UG6B

@misc{pith2026250715628,
  author       = {Pith},
  title        = {Pith review of: A Survey on Efficiency Optimization Techniques for DNN-based Video Analytics: Process Systems, Algorithms, and Applications},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P2I5UG6B}},
  note         = {Machine review of arXiv:2507.15628}
}
read the original abstract

The explosive growth of video data in recent years has brought higher demands for video analytics, where accuracy and efficiency remain the two primary concerns. Deep neural networks (DNNs) have been widely adopted to ensure accuracy; however, improving their efficiency in video analytics remains an open challenge. Different from existing surveys that make summaries of DNN-based video mainly from the accuracy optimization aspect, in this survey, we aim to provide a thorough review of optimization techniques focusing on the improvement of the efficiency of DNNs in video analytics. We organize existing methods in a bottom-up manner, covering multiple perspectives such as hardware support, data processing, operational deployment, etc. Finally, based on the optimization framework and existing works, we analyze and discuss the problems and challenges in the performance optimization of DNN-based video analytics.

Figures

Figures reproduced from arXiv: 2507.15628 by the authors.

Figure 1
Figure 1. Overview of the efficiency optimization work for DNN-based video analytics from the bottom to the [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Overview of computing system optimization. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Edge-Cloud task partition workflow. 3.2.1 Edge-Cloud Data Preprocessing. When processing video data in the cloud, increasing the number of frames analyzed by intensive models raises computation costs and degrades operational efficiency. To address this issue, applying pre-processing steps at the edge, such as edge filtering and video encoding, helps reduce data transmission and alleviate cloud workload. Video Frame … view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Improving DNN algorithm efficiency through spatial and temporal optimization. [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]
Figure 5
Figure 5. Figure 5: DNN-Based spatial dimension optimization focuses on: (a) Pixel-level: adaptive convolution parameters. [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]
Figure 6
Figure 6. Figure 6: DNN-Based temporal dimension optimization focuses on: (a) Feature-wise propagation. Aggregating [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

153 extracted references · 65 canonical work pages

  1. [63]

    Christopher Canel, Thomas Kim, Giulio Zhou, Conglong Li, Hyeontaek Lim, David G Andersen, Michael Kaminsky, and Subramanya Dulloor. 2019. Scaling video analytics on constrained edge nodes. Proceedings of Machine Learning and Systems 1 (2019), 406–417

  2. [34]

    Thomas B Preußer, Giulio Gambardella, Nicholas Fraser, and Michaela Blott. 2018. Inference of quantized neural networks on heterogeneous all-programmable devices. In 2018 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 833–838

  3. [36]

    Ximeng Sun, Rameswar Panda, Chun-Fu Richard Chen, Aude Oliva, Rogerio Feris, and Kate Saenko. 2021. Dynamic network quantization for efficient video inference. In IEEE/CVF ICCV. 7375–7385

  4. [37]

    Jung-Woo Chang, Keon-Woo Kang, and Suk-Ju Kang. 2018. An energy-efficient FPGA-based deconvolutional neural networks accelerator for single image super-resolution. IEEE Transactions on Circuits and Systems for Video Technology 30, 1 (2018), 281–295

  5. [1]

    Minh T Nguyen, Linh H Truong, and Trang TH Le. 2021. Video surveillance processing algorithms utilizing artificial intelligent (AI) for unmanned autonomous vehicles (UAVs). MethodsX 8 (2021), 101472

  6. [2]

    Hassnaa Moustafa, Eve M Schooler, and Jessica McCarthy. 2017. Reverse cdn in fog computing: The lifecycle of video data in connected and autonomous vehicles. In 2017 IEEE FWC. IEEE, 1–5

  7. [3]

    Jiasong Zhu, Ke Sun, Sen Jia, Qingquan Li, Xianxu Hou, Weidong Lin, Bozhi Liu, and Guoping Qiu. 2018. Urban Traffic Density Estimation Based on Ultrahigh-Resolution UAV Video and Deep Neural Network. IEEE J-STARS 11, 12 (2018), 4968–4981. doi:10.1109/JSTARS.2018.2879368

  8. [4]

    Junjue Wang, Ziqiang Feng, Zhuo Chen, Shilpa George, Mihir Bala, Padmanabhan Pillai, Shao-Wen Yang, and Mahadev Satyanarayanan. 2018. Bandwidth-efficient live video analytics for drones via edge computing. In 2018 IEEE/ACM SEC. IEEE, 159–173

Show all 153 references
  1. [5]

    Chang Wen Chen. 2021. Drones as internet of video things front-end sensors: challenges and opportunities. Discover Internet of Things 1, 1 (2021), 1–12

  2. [6]

    Yu Zhao, Yue Yin, and Guan Gui. 2020. Lightweight Deep Learning Based Intelligent Edge Surveillance Techniques. IEEE TCCN 6, 4 (2020), 1146–1154. doi:10.1109/TCCN.2020.2999479

  3. [7]

    Yun-Xia Liu, Yang Yang, Aijun Shi, Peng Jigang, and Liu Haowei. 2019. Intelligent monitoring of indoor surveillance video based on deep learning. In 21st ICACT. 648–653. doi:10.23919/ICACT.2019.8701964

  4. [8]

    Jyotsna, and J

    C.V Amrutha, C. Jyotsna, and J. Amudha. 2020. Deep Learning Approach for Suspicious Activity Detection from Surveillance Video. In 2nd ICIMIA. 335–339. doi:10.1109/ICIMIA48430.2020.9074920

  5. [9]

    GMDT Forecast et al. 2019. Cisco visual networking index: global mobile data traffic forecast update, 2017–2022. Update 2017 (2019), 2022

  6. [10]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. ImageNet Classification with Deep Convolutional Neural Networks. In NeurIPS, F. Pereira, C.J. Burges, L. Bottou, and K.Q. Weinberger (Eds.), Vol. 25. Curran Associates, Inc

  7. [11]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In 2016 IEEE CVPR. 770–778. doi:10.1109/CVPR.2016.90

  8. [12]

    Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)

  9. [13]

    Jian He, Ghufran Baig, and Lili Qiu. 2021. Real-Time Deep Video Analytics on Mobile Devices. In Proceedings of the Twenty-second International Symposium on Theory, Algorithmic Foundations, and Protocol Design for Mobile Networks and Mobile Computing. 81–90

  10. [14]

    Yu Kong and Yun Fu. 2022. Human action recognition and prediction: A survey. IJCV 130, 5 (2022), 1366–1401

  11. [15]

    Sergiu Oprea, Pablo Martinez-Gonzalez, Alberto Garcia-Garcia, John Alejandro Castro-Vargas, Sergio Orts-Escolano, Jose Garcia-Rodriguez, and Antonis Argyros. 2020. A review on deep learning techniques for video prediction. IEEE TPAMI (2020)

  12. [16]

    Qaisar Abbas, Mostafa EA Ibrahim, and M Arfan Jaffar. 2018. Video scene analysis: an overview and challenges on deep learning algorithms. Multimedia Tools and Applications 77, 16 (2018), 20415–20453

  13. [17]

    Nadia Mumtaz, Naveed Ejaz, Shabana Habib, Syed Muhammad Mohsin, Prayag Tiwari, Shahab S Band, and Neeraj Kumar. 2022. An overview of violence detection techniques: current challenges and future directions. Artif Intell. Rev. (2022), 1–26

  14. [18]

    Ratnabali Pal, Arif Ahmed Sekh, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy, and Dilip K Prasad. 2021. Topic-based video analysis: A survey. CSUR 54, 6 (2021), 1–34

  15. [19]

    Arthur G Money and Harry Agius. 2008. Video summarisation: A conceptual framework and survey of the state of the art. JVCIR 19, 2 (2008), 121–143

  16. [20]

    Qingyang Zhang, Hui Sun, Xiaopei Wu, and Hong Zhong. 2019. Edge video analytics for public safety: A review. Proc. IEEE 107, 8 (2019), 1675–1696

  17. [21]

    Huang-Chia Shih. 2017. A survey of content-aware video analysis for sports. IEEE TCSVT 28, 5 (2017), 1212–1231

  18. [22]

    Raz Birman, Yoram Segal, and Ofer Hadar. 2020. Overview of research in the field of video compression using deep neural networks. Multimedia Tools and Applications 79, 17 (2020), 11699–11722

  19. [23]

    Renjie Xu, Saiedeh Razavi, and Rong Zheng. 2022. Deep Learning-Driven Edge Video Analytics: A Survey. arXiv preprint arXiv:2211.15751 (2022)

  20. [24]

    Muhammad Usman Karim Khan, Muhammad Shafique, and Jörg Henkel. 2013. AMBER: Adaptive energy management for on-chip hybrid video memories. In 2013 IEEE/ACM ICCAD. IEEE, 405–412

  21. [25]

    Felipe Sampaio, Muhammad Shafique, Bruno Zatt, Sergio Bampi, and Jörg Henkel. 2014. Energy-efficient architecture for advanced video memory. In 2014 IEEE/ACM ICCAD. IEEE, 132–139

  22. [26]

    Hengyu Zhao, Hongbin Sun, Qiang Yang, Tai Min, and Nanning Zheng. 2016. Exploring the use of volatile STT-RAM for energy efficient video processing. In 17th ISQED. IEEE, 81–87. J. ACM, Vol. 37, No. 4, Article 111. Publication date: August 2025. A Survey on Efficiency Optimizat...

  23. [27]

    Brandon Haynes, Artem Minyaylov, Magdalena Balazinska, Luis Ceze, and Alvin Cheung. 2017. Visualcloud demon- stration: A dbms for virtual reality. In Int. Conf. Manage. Data (SIGMOD) . 1615–1618

  24. [28]

    Brandon Haynes, Amrita Mazumdar, Magdalena Balazinska, Luis Ceze, and Alvin Cheung. 2018. Lightdb: A dbms for virtual reality video. Proceedings of the VLDB Endowment 11, 10 (2018)

  25. [29]

    Amrita Mazumdar, Brandon Haynes, Magda Balazinska, Luis Ceze, Alvin Cheung, and Mark Oskin. 2019. Perceptual compression for video storage and processing systems. In Proceedings of the ACM Symposium on Cloud Computing . 179–192

  26. [30]

    Tiantu Xu, Luis Materon Botelho, and Felix Xiaozhu Lin. 2019. Vstore: A data store for analytics on large videos. In Proceedings of the Fourteenth EuroSys Conference 2019 . 1–17

  27. [31]

    Brandon Haynes, Maureen Daum, Dong He, Amrita Mazumdar, Magdalena Balazinska, Alvin Cheung, and Luis Ceze

  28. [32]

    Maureen Daum, Brandon Haynes, Dong He, Amrita Mazumdar, and Magdalena Balazinska. 2021. TASM: A tile-based storage manager for video analytics. In 2021 IEEE ICDE. IEEE, 1775–1786

  29. [33]

    Weisong Shi, Jie Cao, Quan Zhang, Youhuizi Li, and Lanyu Xu. 2016. Edge computing: Vision and challenges. IEEE Internet Things J. 3, 5 (2016), 637–646

  30. [35]

    Matthieu Courbariaux, Itay Hubara, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. 2016. Binarized neural networks: Training deep neural networks with weights and activations constrained to+ 1 or-1. arXiv preprint arXiv:1602.02830 (2016)

  31. [38]

    Song Han, Huizi Mao, and William J Dally. 2019. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv 2015. arXiv preprint arXiv:1510.00149 (2019)

  32. [39]

    Chun-Ya Tsai, De-Qin Gao, and Shanq-Jang Ruan. 2020. An effective hybrid pruning architecture of dynamic convolution for surveillance videos. Journal of Visual Communication and Image Representation 70 (2020), 102798

  33. [40]

    Zhenzhen Wang, Weixiang Hong, Yap-Peng Tan, and Junsong Yuan. 2019. Pruning 3d filters for accelerating 3d convnets. IEEE TMM 22, 8 (2019), 2126–2137

  34. [41]

    Cheng Dai, Xingang Liu, Hao Xu, Laurence Tianruo Yang, and M Jamal Deen. 2021. Hybrid deep model for human behavior understanding on industrial internet of video things. IEEE TII 18, 10 (2021), 7000–7008

  35. [42]

    Yuan Cheng, Guangya Li, Ngai Wong, Hai-Bao Chen, and Hao Yu. 2020. DEEPEYE: A deeply tensor-compressed neural network for video comprehension on terminal devices. ACM TECS 19, 3 (2020), 1–25

  36. [43]

    Cheng Dai, Xingang Liu, Hongqiang Cheng, Laurence T Yang, and M Jamal Deen. 2021. Compressing deep model with pruning and tucker decomposition for smart embedded systems. IEEE Internet Things J. 9, 16 (2021), 14490–14500

  37. [44]

    Ravi Teja Mullapudi, Steven Chen, Keyi Zhang, Deva Ramanan, and Kayvon Fatahalian. 2019. Online model distillation for efficient video inference. In IEEE/CVF ICCV. 3573–3582

  38. [45]

    Roy Miles, Mehmet Kerim Yucel, Bruno Manganelli, and Albert Saa-Garriga. 2023. MobileVOS: Real-Time Video Object Segmentation Contrastive Learning meets Knowledge Distillation. arXiv preprint arXiv:2303.07815 (2023)

  39. [46]

    Duc-Quang Vu, Ngan TH Le, and Jia-Ching Wang. 2022. (2+ 1) D Distilled ShuffleNet: A Lightweight Unsupervised Distillation Network for Human Action Recognition. In 26th ICPR. IEEE, 3197–3203

  40. [47]

    Ayesha Younis, Li Shixin, Shelembi Jn, and Zhang Hai. 2020. Real-time object detection using pre-trained deep learning models MobileNet-SSD. In ICCDE. 44–48

  41. [48]

    Hao Zhang, Shenqi Lai, Yaxiong Wang, Zongyang Da, Yujie Dun, and Xueming Qian. 2023. Scgnet: Shifting and cascaded group network. IEEE TCSVT (2023)

  42. [49]

    Ran Xu, Chen-lin Zhang, Pengcheng Wang, Jayoung Lee, Subrata Mitra, Somali Chaterji, Yin Li, and Saurabh Bagchi

  43. [50]

    Jayoung Lee, Pengcheng Wang, Ran Xu, Venkat Dasari, Noah Weston, Yin Li, Saurabh Bagchi, and Somali Chaterji

  44. [51]

    Ran Xu, Fangzhou Mu, Jayoung Lee, Preeti Mukherjee, Somali Chaterji, Saurabh Bagchi, and Yin Li. 2022. SMAR- TADAPT: Multi-branch Object Detection Framework for Videos on Mobiles. In IEEE/CVF CVPR. 2528–2538. J. ACM, Vol. 37, No. 4, Article 111. Publication date: August 2025. ...

  45. [52]

    Ziyi Wang, Rui Zhang, Songyu Zhang, Wei Cheng, Wendong Wang, and Yong Cui. 2025. Edge-assisted adaptive configuration for serverless-based video analytics. IEEE Transactions on Networking (2025)

  46. [53]

    In 5th EMDL

    Benchmarking video object detection systems on embedded devices under resource contention. In 5th EMDL. 19–24

  47. [54]

    Xiao Zeng, Biyi Fang, Haichen Shen, and Mi Zhang. 2020. Distream: scaling live video analytics with workload- adaptive distributed edge intelligence. In Proceedings of the 18th Conference on Embedded Networked Sensor Systems . 409–421

  48. [55]

    Wuyang Zhang, Zhezhi He, Luyang Liu, Zhenhua Jia, Yunxin Liu, Marco Gruteser, Dipankar Raychaudhuri, and Yanyong Zhang. 2021. Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloading. In 27th MobiCom. 201–214

  49. [56]

    Hui Sun, Ying Yu, Kewei Sha, and Hong Zhong. 2022. EdgeEye: A Data-Driven Approach for Optimal Deployment of Edge Video Analytics. IEEE Internet Things J. 9, 19 (2022), 19273–19295

  50. [57]

    Angela H Jiang, Daniel L-K Wong, Christopher Canel, Lilia Tang, Ishan Misra, Michael Kaminsky, Michael A Kozuch, Padmanabhan Pillai, David G Andersen, and Gregory R Ganger. 2018. Mainstream: Dynamic{Stem-Sharing} for {Multi-Tenant} Video Processing. In 2018 USENIX Annual Techn...

  51. [58]

    Biyi Fang, Xiao Zeng, and Mi Zhang. 2018. Nestdnn: Resource-aware multi-tenant on-device deep learning for continuous mobile vision. In 24th MobiCom. 115–127

  52. [59]

    Daniel Zhang, Yue Ma, Chao Zheng, Yang Zhang, X Sharon Hu, and Dong Wang. 2018. Cooperative-competitive task allocation in edge computing for delay-sensitive social sensing. In 2018 IEEE/ACM SEC. IEEE, 243–259

  53. [60]

    Qian Li, Jiayu Song, Jiangbo Ning, and Jianping Yuan. 2019. The Detailed Data on the Neural Compute Stick Acceleration Performance. In 2019 Chinese Automation Congress (CAC). IEEE, 4959–4962

  54. [61]

    Oleksandra Aleksandrova and Yevgen Bashkov. 2020. Face recognition systems based on Neural Compute Stick 2, CPU, GPU comparison. In 2020 IEEE ATIT. IEEE, 104–107

  55. [62]

    [n. d.]. Neural compute Stick Documentation. https://software.intel.com/enus/movidius-ncs

  56. [64]

    Yuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Yufei Wang, Guoqing Harry Xu, and Ravi Netravali. 2020. Reducto: On-camera filtering for resource-efficient real-time video analytics. In ACM SIGCOMM Conference. 359–376

  57. [65]

    [n. d.]. Google LLC. https://coral.withgoogle.com/tutorialgoogls/accelerator-datasheet Google Coral Datasheet

  58. [66]

    2018.{VideoChef}: Efficient Approximation for Streaming Video Processing Pipelines

    Ran Xu, Jinkyu Koo, Rakesh Kumar, Peter Bai, Subrata Mitra, Sasa Misailovic, and Saurabh Bagchi. 2018.{VideoChef}: Efficient Approximation for Streaming Video Processing Pipelines. In 2018 USENIX Annual Technical Conference (USENIX ATC 18). 43–56

  59. [67]

    Kuntai Du, Qizheng Zhang, Anton Arapin, Haodong Wang, Zhengxu Xia, and Junchen Jiang. 2022. AccMPEG: Optimizing Video Encoding for Video Analytics. arXiv preprint arXiv:2204.12534 (2022)

  60. [68]

    Jude Tchaye-Kondi, Yanlong Zhai, Jun Shen, Dong Lu, and Liehuang Zhu. 2022. Smartfilter: an edge system for real-time application-guided video frames filtering. IEEE Internet Things J. 9, 23 (2022), 23772–23785

  61. [69]

    Sandeep P Chinchali, Eyal Cidon, Evgenya Pergament, Tianshu Chu, and Sachin Katti. 2018. Neural networks meet physical networks: Distributed inference between edge devices and the cloud. InProceedings of the 17th ACM Workshop on Hot Topics in Networks . 50–56

  62. [70]

    Ben Zhang, Xin Jin, Sylvia Ratnasamy, John Wawrzynek, and Edward A Lee. 2018. Awstream: Adaptive wide-area streaming analytics. In ACM SIGCOMM. 236–252

  63. [71]

    Kuntai Du, Ahsan Pervaiz, Xin Yuan, Aakanksha Chowdhery, Qizheng Zhang, Henry Hoffmann, and Junchen Jiang

  64. [72]

    In ACM SIGCOMM

    Server-driven video streaming for deep learning inference. In ACM SIGCOMM. 557–570

  65. [73]

    Zhuoran Song, Chunyu Qi, Fangxin Liu, Naifeng Jing, and Xiaoyao Liang. 2024. Cmc: Video transformer acceleration via codec assisted matrix condensing. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Syste...

  66. [74]

    Hui Sun, Ying Yu, Kewei Sha, and Bendong Lou. 2019. mVideo: edge computing based mobile video processing systems. IEEE Access 8 (2019), 11615–11623

  67. [75]

    Yiding Wang, Weiyan Wang, Junxue Zhang, Junchen Jiang, and Kai Chen. 2019. Bridging the{Edge-Cloud} Barrier for Real-time Advanced Vision Analytics. In 11th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 19)

  68. [76]

    Huaizheng Zhang, Meng Shen, Yizheng Huang, Yonggang Wen, Yong Luo, Guanyu Gao, and Kyle Guan. 2021. A serverless cloud-fog platform for dnn-based video analytics with incremental learning.arXiv preprint arXiv:2102.03012 (2021)

  69. [77]

    Yiping Kang, Johann Hauswald, Cao Gao, Austin Rovinski, Trevor Mudge, Jason Mars, and Lingjia Tang. 2017. Neurosurgeon: Collaborative intelligence between the cloud and mobile edge. ACM SIGARCH Computer Architecture News 45, 1 (2017), 615–629

  70. [78]

    Chuang Hu, Wei Bao, Dan Wang, and Fengming Liu. 2019. Dynamic adaptive DNN surgery for inference acceleration on the edge. In IEEE INFOCOM 2019-IEEE Conference on Computer Communications . IEEE, 1423–1431

  71. [79]

    Swarnava Dey, Arijit Mukherjee, and Arpan Pal. 2019. Embedded deep inference in practice: Case for model partitioning. In Proceedings of the 1st Workshop on Machine Learning on Edge in Sensor Systems . 25–30

  72. [80]

    Peiyin Xing, Xiaofei Liu, Peixi Peng, Tiejun Huang, and Yonghong Tian. 2021. Allocating DNN Layers Computation Between Front-End Devices and The Cloud Server for Video Big Data Processing. In 2021-2021 IEEE ICASSP. IEEE, 8483–8487. J. ACM, Vol. 37, No. 4, Article 111. Publicat...

  73. [81]

    Xukan Ran, Haolianz Chen, Xiaodan Zhu, Zhenming Liu, and Jiasi Chen. 2018. Deepdecision: A mobile deep learning framework for edge video analytics. In IEEE INFOCOM 2018-IEEE Conference on Computer Communications . IEEE, 1421–1429

  74. [82]

    Haichen Shen, Lequn Chen, Yuchen Jin, Liangyu Zhao, Bingyu Kong, Matthai Philipose, Arvind Krishnamurthy, and Ravi Sundaram. 2019. Nexus: A GPU cluster engine for accelerating DNN-based video analysis. In 27th ACM SOSP. 322–337

  75. [83]

    Ke-Jou Hsu, Ketan Bhardwaj, and Ada Gavrilovska. 2019. Couper: Dnn model slicing for visual analytics containers at the edge. In Proceedings of the 4th ACM/IEEE SEC . 179–194

  76. [84]

    Seungyeop Han, Haichen Shen, Matthai Philipose, Sharad Agarwal, Alec Wolman, and Arvind Krishnamurthy. 2016. Mcdnn: An approximation-based execution framework for deep stream processing under resource constraints. In 14th ACM MobiSys. 123–136

  77. [85]

    Cohen, and Babak Ehteshami Bejnordi

    Amirhossein Habibian, Davide Abati, Taco S. Cohen, and Babak Ehteshami Bejnordi. 2021. Skip-Convolutions for Efficient Video Processing. arXiv:2104.11487 [cs.CV]

  78. [86]

    Bo Li, Junjie Yan, Wei Wu, Zheng Zhu, and Xiaolin Hu. 2018. High performance visual tracking with siamese region proposal network. In IEEE CVPR. 8971–8980

  79. [87]

    Haitao Zhang, Bingchang Tang, Xin Geng, and Huadong Ma. 2018. Learning driven parallelization for large-scale video workload in hybrid CPU-GPU cluster. In 47th ICPP. 1–10

  80. [88]

    Hang Su, Varun Jampani, Deqing Sun, Orazio Gallo, Erik Learned-Miller, and Jan Kautz. 2019. Pixel-adaptive convolutional neural networks. In IEEE/CVF CVPR. 11166–11175

  81. [89]

    Yuning Chai. 2019. Patchwork: A patch-wise attention network for efficient object detection and segmentation in video streams. In IEEE/CVF ICCV. 3415–3424

  82. [90]

    Ting-Wu Chin, Ruizhou Ding, and Diana Marculescu. 2019. Adascale: Towards real-time video object detection using adaptive scaling. Proceedings of Machine Learning and Systems 1 (2019), 431–441

  83. [91]

    Huizi Mao, Sibo Zhu, Song Han, and William J Dally. 2021. PatchNet–Short-range Template Matching for Efficient Video Processing. arXiv preprint arXiv:2103.07371 (2021)

  84. [92]

    Yulin Wang, Zhaoxi Chen, Haojun Jiang, Shiji Song, Yizeng Han, and Gao Huang. 2021. Adaptive focus for efficient video recognition. In IEEE/CVF ICCV. 16249–16258

  85. [93]

    Chaoxu Guo, Bin Fan, Jie Gu, Qian Zhang, Shiming Xiang, Veronique Prinet, and Chunhong Pan. 2019. Progressive sparse local attention for video object detection. In Proceedings of the IEEE/CVF ICCV . 3909–3918

  86. [94]

    Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. 2022. Video swin transformer. In IEEE/CVF CVPR. 3202–3211

  87. [95]

    Yue Meng, Chung-Ching Lin, Rameswar Panda, Prasanna Sattigeri, Leonid Karlinsky, Aude Oliva, Kate Saenko, and Rogerio Feris. 2020. Ar-net: Adaptive frame resolution for efficient action recognition. In 16th ECCV Part VII. Springer, 86–104

  88. [96]

    Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. 2018. Non-local neural networks. In IEEE CVPR. 7794–7803

  89. [97]

    Bruno Korbar, Du Tran, and Lorenzo Torresani. 2019. Scsampler: Sampling salient clips from video for efficient action recognition. In Proceedings of the IEEE/CVF ICCV . 6232–6242

  90. [98]

    Zuxuan Wu, Caiming Xiong, Chih-Yao Ma, Richard Socher, and Larry S Davis. 2019. Adaframe: Adaptive frame selection for fast video recognition. In Proceedings of the IEEE/CVF CVPR . 1278–1287

  91. [99]

    Mykhailo Shvets, Wei Liu, and Alexander C Berg. 2019. Leveraging long-range temporal relationships between proposals for video object detection. In IEEE/CVF ICCV. 9756–9764

  92. [100]

    Tao Gong, Kai Chen, Xinjiang Wang, Qi Chu, Feng Zhu, Dahua Lin, Nenghai Yu, and Huamin Feng. 2021. Temporal ROI align for video object recognition. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 1442–1450

  93. [101]

    Chun-Han Yao, Chen Fang, Xiaohui Shen, Yangyue Wan, and Ming-Hsuan Yang. 2020. Video object detection via object-level temporal aggregation. In 16th ECCV Part XIV . Springer, 160–177

  94. [102]

    Jintao Lin, Haodong Duan, Kai Chen, Dahua Lin, and Limin Wang. 2022. Ocsampler: Compressing videos to one clip with single-step sampling. In IEEE/CVF CVPR. 13894–13903. J. ACM, Vol. 37, No. 4, Article 111. Publication date: August 2025. 111:24 Trovato et al

  95. [103]

    Wenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen, and Shilei Wen. 2019. Multi-agent reinforcement learning based frame sampling for effective untrimmed video recognition. In IEEE/CVF ICCV. 6222–6231

  96. [104]

    Mason Liu, Menglong Zhu, Marie White, Yinxiao Li, and Dmitry Kalenichenko. 2019. Looking fast and slow: Memory-guided mobile video object detection. arXiv preprint arXiv:1903.10172 (2019)

  97. [105]

    Hehe Fan, Zhongwen Xu, Linchao Zhu, Chenggang Yan, Jianjun Ge, and Yi Yang. 2018. Watching a small portion could be as good as watching all: Towards efficient video classification. In IJCAI

  98. [106]

    Wenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen, Yi Yang, and Shilei Wen. 2020. Dynamic inference: A new approach toward efficient video action recognition. In IEEE/CVF CVPRW. 676–677

  99. [107]

    Hanul Kim, Mihir Jain, Jun-Tae Lee, Sungrack Yun, and Fatih Porikli. 2021. Efficient action recognition via dynamic knowledge propagation. In IEEE/CVF ICCV. 13719–13728

  100. [108]

    Yanni Yang, Huansheng Song, Shijie Sun, Yan Chen, Xinyao Tang, and Qin Shi. 2023. A feature temporal attention based interleaved network for fast video object detection. Journal of Ambient Intelligence and Humanized Computing 14, 1 (2023), 497–509

  101. [109]

    Hao Fu, Shanjiang Tang, Ce Yu, Yusen Li, Jizhou Sun, and Yanjie Liu. 2021. DVQShare: An Analytics System for DNN-based Video Queries. In 2021 IEEE/ACM 21st CCGrid. IEEE, 166–175

  102. [110]

    Haoyu Zhang, Ganesh Ananthanarayanan, Peter Bodik, Matthai Philipose, Paramvir Bahl, and Michael J Freedman

  103. [111]

    Amir Ghodrati, Babak Ehteshami Bejnordi, and Amirhossein Habibian. 2021. Frameexit: Conditional early exiting for efficient video recognition. In IEEE/CVF CVPR. 15608–15618

  104. [112]

    Ziliang Lai, Chenxia Han, Chris Liu, Pengfei Zhang, Eric Lo, and Ben Kao. 2021. Top-K Deep Video Analytics: A Probabilistic Approach. In Int. Conf. Manage. Data (SIGMOD) . 1037–1050

  105. [113]

    Daniel Y Fu, Will Crichton, James Hong, Xinwei Yao, Haotian Zhang, Anh Truong, Avanika Narayan, Maneesh Agrawala, Christopher Ré, and Kayvon Fatahalian. 2019. Rekall: Specifying video events using compositions of spatiotemporal labels. arXiv preprint arXiv:1910.02993 (2019)

  106. [114]

    Francisco Romero, Johann Hauswald, Aditi Partap, Daniel Kang, Matei Zaharia, and Christos Kozyrakis. 2022. Optimizing Video Analytics with Declarative Model Relationships. Proceedings of the VLDB Endowment 16, 3 (2022), 447–460

  107. [115]

    Takashi Isobe, Fang Zhu, Xu Jia, and Shengjin Wang. 2020. Revisiting temporal modeling for video super-resolution. arXiv preprint arXiv:2008.05765 (2020)

  108. [116]

    Daniel Kang, John Emmons, Firas Abuzaid, Peter Bailis, and Matei Zaharia. 2017. Noscope: optimizing neural network queries over video at scale. arXiv preprint arXiv:1703.02529 (2017)

  109. [117]

    Daniel Kang, Peter Bailis, and Matei Zaharia. 2018. Blazeit: Optimizing declarative aggregation and limit queries for neural network-based video analytics. arXiv preprint arXiv:1805.01046 (2018)

  110. [118]

    Dario Fuoli, Martin Danelljan, Radu Timofte, and Luc Van Gool. 2023. Fast online video super-resolution with deformable attention pyramid. In IEEE/CVF W ACV. 1735–1744

  111. [119]

    Hoang Nguyen Ngoc, Nhat Nguyen Xuan, Trung H Bui, Dao Huu Hung, Steven QH Truong, and Vu Hoang. 2023. An efficient approach for real-time abnormal human behavior recognition on surveillance cameras. In 2023 IEEE FG. IEEE, 1–6

  112. [120]

    Wenguan Wang, Jianbing Shen, and Ling Shao. 2017. Video salient object detection via fully convolutional networks. IEEE TIP 27, 1 (2017), 38–49

  113. [121]

    Dario Fuoli, Shuhang Gu, and Radu Timofte. 2019. Efficient video super-resolution through recurrent latent space propagation. In 2019 IEEE/CVF ICCVW. IEEE, 3476–3485

  114. [122]

    Takashi Isobe, Xu Jia, Xin Tao, Changlin Li, Ruihuang Li, Yongjie Shi, Jing Mu, Huchuan Lu, and Yu-Wing Tai

  115. [123]

    Mengyang Liu, Anran Tang, Huitian Wang, Lin Shen, Yunhan Chang, Guangxing Cai, Daheng Yin, Fang Dong, and Wei Zhao. 2022. Accelerating Multi-Object Tracking in Edge Computing Environment with Time-Spatial Optimization. In 9th Int. Conf. Adv. Cloud Big Data (CBD) . IEEE, 279–284

  116. [124]

    Favyen Bastani, Songtao He, Arjun Balasingam, Karthik Gopalakrishnan, Mohammad Alizadeh, Hari Balakrishnan, Michael Cafarella, Tim Kraska, and Sam Madden. 2020. Miris: Fast object track queries in video. In Int. Conf. Manage. Data (SIGMOD). 1907–1921

  117. [125]

    Fabio Tosi, Filippo Aleotti, Pierluigi Zama Ramirez, Matteo Poggi, Samuele Salti, Luigi Di Stefano, and Stefano Mattoccia. 2020. Distilled semantics for comprehensive scene understanding from videos. In IEEE/CVF CVPR. 4654– 4665

  118. [126]

    Yuan Cheng, Yuchao Yang, Hai-Bao Chen, Ngai Wong, and Hao Yu. 2021. S3-net: a fast and lightweight video scene understanding network by single-shot segmentation. In IEEE/CVF W ACV. 3329–3337. J. ACM, Vol. 37, No. 4, Article 111. Publication date: August 2025. A Survey on Effic...

  119. [127]

    Fang Guo, Wenguan Wang, Ziyi Shen, Jianbing Shen, Ling Shao, and Dacheng Tao. 2019. Motion-aware rapid video saliency detection. IEEE TCSVT 30, 12 (2019), 4887–4898

  120. [128]

    Feng Liu, Zhigang Han, Qian Li, and Caihui Cui. 2021. Dynamic Target Tracking and Geospatial Transformation in Place-Based on DNN. In Geoinformatics. IEEE, 1–5

  121. [129]

    Khan Muhammad, Tanveer Hussain, Javier Del Ser, Vasile Palade, and Victor Hugo C De Albuquerque. 2019. DeepReS: A deep learning-based video summarization strategy for resource-constrained industrial surveillance scenarios. IEEE TII 16, 9 (2019), 5938–5947

  122. [130]

    Muhammad Ajmal, Mudasser Naseer, Farooq Ahmad, and Asma Saleem. 2017. Human motion trajectory analysis based video summarization. In 2017 16th IEEE ICMLA . IEEE, 550–555

  123. [131]

    Gionatan Gallo, Francesco Di Rienzo, Pietro Ducange, Vincenzo Ferrari, Alessandro Tognetti, and Carlo Vallati. 2021. A smart system for personal protective equipment detection in industrial environments based on deep learning. In 2021 IEEE SMARTCOMP. IEEE, 222–227

  124. [132]

    Nipun D Nath, Amir H Behzadan, and Stephanie G Paal. 2020. Deep learning for site safety: Real-time detection of personal protective equipment. Automation in Construction 112 (2020), 103085

  125. [133]

    Huaizheng Zhang, Yuanming Li, Qiming Ai, Yong Luo, Yonggang Wen, Yichao Jin, and Nguyen Binh Duong Ta. 2020. Hysia: Serving DNN-Based Video-to-Retail Applications in Cloud. In 28th ACM MM. 4457–4460

  126. [134]

    Florian Vandecasteele, Karel Vandenbroucke, Dimitri Schuurman, and Steven Verstockt. 2017. Spott: On-the-spot e-commerce for television using deep learning-based video analysis techniques. ACM TOMM 13, 3s (2017), 1–16

  127. [135]

    Huawei Launches the Atlas Intelligent Computing Platform to Fuel an AI Future with Supreme Compute Power

    2018. Huawei Launches the Atlas Intelligent Computing Platform to Fuel an AI Future with Supreme Compute Power. https://www.huawei.com/en/news/2018/10/atlas-intelligent-computing-platform. [Online; accessed 28 Oct. 2021]

  128. [136]

    Sophon Edge

    2020. Sophon Edge. https://www.96boards.org/product/sophon-edge. [Online; accessed 29 Dec. 2020]

  129. [137]

    Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al. 2022. InternVideo: General Video Foundation Models via Generative and Discriminative Learning. arXiv preprint arXiv:2212.03191 (2022)

  130. [138]

    Junke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo, Luowei Zhou, Yucheng Zhao, Yujia Xie, Ce Liu, Yu-Gang Jiang, and Lu Yuan. 2022. Omnivl: One foundation model for image-language and video-language tasks.arXiv preprint arXiv:2209.07526 (2022)

  131. [139]

    Norman P Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, et al. 2023. Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings. arXiv prepri...

  132. [140]

    CUDA NVIDIA. 2020. Toolkit 8.0 download (November 2019): https://developer. nvidia. com/cuda-toolkit-archive . Technical Report. Accesed 28/01

  133. [141]

    Kunchang Li, Yali Wang, Yizhuo Li, Yi Wang, Yinan He, Limin Wang, and Yu Qiao. 2023. Unmasked teacher: Towards training-efficient video foundation models. arXiv preprint arXiv:2303.16058 (2023)

  134. [142]

    Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258 (2021)

  135. [143]

    Luyi Chang, Zhe Zhang, Pei Li, Shan Xi, Wei Guo, Yukang Shen, Zehui Xiong, Jiawen Kang, Dusit Niyato, Xiuquan Qiao, et al. 2022. 6G-enabled edge AI for Metaverse: Challenges, methods, and future research directions. Journal of Communications and Information Networks 7, 2 (2022...

  136. [144]

    Yiding Wang, Weiyan Wang, Duowen Liu, Xin Jin, Junchen Jiang, and Kai Chen. 2022. Enabling edge-cloud video analytics for robotics applications. IEEE TCC (2022)

  137. [145]

    Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao. 2023. Videomae v2: Scaling video masked autoencoders with dual masking. arXiv preprint arXiv:2303.16727 (2023)

  138. [146]

    Ce Zhou, Qian Li, Chen Li, Jun Yu, Yixin Liu, Guangjing Wang, Kai Zhang, Cheng Ji, Qiben Yan, Lifang He, et al

  139. [152]

    Muhammad Shoaib Siddiqui, Toqeer Ali Syed, Adnan Nadeem, Waqas Nawaz, and Ahmad Alkhodre. 2022. Virtual Tourism and Digital Heritage: An Analysis of VR/AR Technologies and Applications. IJACSA 13, 7 (2022)

  140. [153]

    Abid Yaqoob, Ting Bi, and Gabriel-Miro Muntean. 2020. A survey on adaptive 360 video streaming: Solutions, challenges and opportunities. IEEE Communications Surveys & Tutorials 22, 4 (2020), 2801–2838. Received 12 June 2025; revised –; accepted – J. ACM, Vol. 37, No. 4, Articl...

  141. [2017]

    In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17)

    Live Video Analytics at Scale with Approximation and {Delay-Tolerance}. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17) . 377–392

  142. [2020]

    In Proceedings of the 18th Conference on Embedded Networked Sensor Systems

    ApproxDet: content and contention-aware approximate object detection for mobiles. In Proceedings of the 18th Conference on Embedded Networked Sensor Systems . 449–462

  143. [2021]

    Vss: A storage system for video analytics. In Int. Conf. Manage. Data (SIGMOD) . 685–696

  144. [2022]

    In IEEE/CVF CVPR

    Look back and forth: Video super-resolution with explicit temporal difference modeling. In IEEE/CVF CVPR. 17411–17420

  145. [2023]

    arXiv preprint arXiv:2302.09419 (2023)

    A comprehensive survey on pretrained foundation models: A history from bert to chatgpt. arXiv preprint arXiv:2302.09419 (2023)

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.