Pith. sign in

REVIEW 3 major objections 2 minor 1 cited by

SHeRL-FL: When Representation Learning Meets Split Learning in Hierarchical Federated Learning

T0 review · 3 major / 2 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read The manuscript pairs a SHeRL-FL abstract with a TinyML survey body.

desk verdict The abstract and full text are two different papers—SHeRL-FL's claimed results are entirely absent from a competent but unrelated TinyML OD survey. read the letter →

arxiv 2508.08339 v1 pith:74P3OX7X submitted 2025-08-11 cs.LG

classification cs.LG
keywords TinyMLobjectdetectionmodelcompressionquantizationpruningknowledgedistillationneuralarchitecturesearchsplitlearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The submission as received has two mismatched layers. The metadata and abstract claim a new framework, SHeRL-FL, that combines split learning and hierarchical federated learning with representation learning at intermediate layers, and report that it cuts data transmission by more than 90% versus centralized FL and HierFL and by 50% versus SplitFed, with experiments on CIFAR-10/100, HAM10000, and ISIC-2018. The full text, however, is a different paper: a survey titled 'Designing Object Detection Models for TinyML' that analyzes quantization, pruning, knowledge distillation, and neural architecture search for object detection on microcontrollers, with comparison tables of MCU-deployed detectors. Read sympathetically, the full-text authors are trying to establish that existing surveys overlook optimization challenges specific to deploying object detectors on TinyML devices, and that a structured map of these techniques, plus benchmark numbers, fills that gap. If the abstract's SHeRL-FL claim were true, it would be a substantial communication-efficiency gain for hierarchical split learning; but the body contains no SHeRL-FL algorithm, no such experiments, and no transmission measurements.

What carries the argument

For the abstract's claimed framework, the key mechanism is representation learning at intermediate layers: clients and edge servers would compute training objectives independently of the cloud, reducing coordination complexity and cutting the volume of data crossing tiers. For the full-text survey, the carrying mechanism is a taxonomy that splits model optimization into parameter removal (pruning), parameter quantization (QAT/PTQ/BNNs), parameter search (NAS), and knowledge transfer (KD), all applied to the backbone–neck–head pipeline of object detectors. The load-bearing evidence is Table 8, which compares MCU-deployed detectors by parameters, MMACs, peak SRAM, and mAP on PASCAL VOC; the ta

What would settle it

A reader can settle the mismatch immediately by searching the full text for 'SHeRL-FL'; it appears nowhere in the body, and no algorithm, no CIFAR/HAM10000/ISIC experiment, and no transmission-volume table is reported, so there is nothing in the manuscript that could confirm or refute the abstract's numbers.

Watch

Extended reading notes

Core claim

The full-text authors' central claim is that previous surveys of lightweight object detection focus on backbones or general edge AI and miss the optimization of detection models under TinyML memory budgets. To close that gap, the survey organizes the field into four compression families—quantization (QAT, PTQ, binary networks), pruning (unstructured, structured, semi-structured), knowledge distillation (feature, multi-teacher, multi-modal, self-, weakly supervised), and neural architecture search (RL-, evolutionary-, gradient-, and hardware-aware)—and compares MCU-optimized detectors on PASCAL VOC, reaching 51.4–74.9% mAP with 53–511 kB peak SRAM. The body does not contain the SHeRL-FL frame

Load-bearing premise

The load-bearing premise for the abstract's central claim is that the manuscript actually contains the SHeRL-FL method and its experiments; the full text instead is a different survey, so the claimed 90% and 50% transmission reductions rest on a document that is not present.

Editorial extensions

If this is right

  • The survey's taxonomy gives TinyML practitioners a direct way to match a compression family (e.g., quantization or NAS) to a specific hardware constraint such as peak SRAM or MMACs.
  • The tabulated MCU detectors show a current operating band of roughly 51–75% mAP on PASCAL VOC at under 800 MMACs, which frames how much accuracy headroom remains for extreme low-power object detection.
  • The open-challenges section singles out energy-efficient SNN-based detectors, high-resolution input handling, and transformer-based architectures as the next targets for TinyML object detection.
  • If the abstract's claimed 90% and 50% transmission reductions were replicated, hierarchical split learning would become a practical bandwidth-saving option for federated training with heterogeneous edge clients.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The claimed 50% cut versus SplitFed suggests the bottleneck it addresses is not body computation but the cross-tier transfer of intermediate activations; a natural follow-up experiment would measure per-round transmission as the cut layer moves through the network.
  • The survey's benchmark numbers come from different papers with different training setups, so a fair comparison of MCU detectors would require re-running the same models under a common training and quantization pipeline.
  • The taxonomy's emphasis on hardware-aware NAS and co-design suggests that future TinyML object detectors will increasingly be optimized jointly with the inference engine and memory scheduler, not just the network weights.
  • Treating the submitted file as two separate documents—a SHeRL-FL systems paper and a TinyML survey—would let each be evaluated on its own evidence.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The submission, as indexed, claims to propose SHeRL-FL, a method integrating split learning and hierarchical federated learning with representation learning, and reports communication reductions of over 90% versus centralized FL/HierFL and 50% versus SplitFed, with experiments on CIFAR-10, CIFAR-100, HAM10000, and ISIC-2018. The full text, however, is not that paper. It is a survey titled "Designing Object Detection Models for TinyML: Foundations, Comparative Analysis, Challenges, and Emerging Solutions" by different authors, covering quantization, pruning, knowledge distillation, and neural architecture search for object detection on microcontrollers. The body contains no mention of SHeRL-FL, split learning, hierarchical FL, or any of the claimed datasets. The central contribution and experimental evidence promised in the abstract are therefore absent from the manuscript.

Significance. If the abstract's claims were supported by a real method and experiments, SHeRL-FL would be a potentially significant communication-efficiency contribution to federated and split learning. However, as submitted, no such method, derivation, or experiment exists in the manuscript, so the significance of the claimed contribution cannot be assessed. Taken on its own terms as a TinyML object-detection survey, the full text has some merits: it offers a structured taxonomy of four optimization families, tabulated model comparisons (Tables 5, 6, 8), and a public repository link. These are useful survey elements, but they do not constitute the federated-learning paper described by the abstract.

major comments (3)
  1. [Abstract; Sections 1 and 9] The abstract promises a method called SHeRL-FL with quantitative results: 'reduces data transmission by over 90% compared to centralized FL and HierFL, and by 50% compared to SplitFed', based on experiments on CIFAR-10, CIFAR-100, HAM10000, and ISIC-2018. The full text is a different paper with a different title and author list: a TinyML object-detection survey. I could not find the term 'SHeRL-FL' anywhere in the body, nor any discussion of split learning, hierarchical federated learning, or the claimed datasets. The central claim of the submission is therefore unsupported by any content in the manuscript. This is a load-bearing mismatch: it is not a local error but the absence of the paper itself.
  2. [Sections 2-9] There is no algorithmic description, training setup, aggregation rule, loss function, baseline configuration, or measured transmission-cost analysis for SHeRL-FL. The body instead surveys quantization, pruning, knowledge distillation, and NAS for object detection. Because the method and experiments are absent, the abstract's assertions about reduced coordination complexity and communication overhead cannot be derived, reproduced, or checked. This is not a gap that can be fixed by adding a missing section; it requires the actual SHeRL-FL manuscript.
  3. [Section 1] Even if the TinyML survey is considered the intended submission, the stated research gap is disjoint from the abstract's claimed contribution. Section 1 defines the gap as the lack of surveys covering optimization techniques for OD on resource-constrained devices, while the abstract frames the contribution as a new FL/SL method. A paper cannot simultaneously be a novel federated-learning algorithm and a survey of TinyML object-detection compression without any connection between the two. The title, abstract, and full text describe different papers.
minor comments (2)
  1. [Section 1] The text contains several typographical errors, e.g., '150,55 billion' (should be '150.55 billion'), 'compartive', and later 'qantization' and 'improvment'. These should be corrected in any revision.
  2. [Fig. 3; Section 6.4.4] The taxonomy figure is dense and the subcategories are not numbered in the figure, making it hard to map to the text. Also, the self-citation [164] in Section 6.4.4 is used as an illustrative HNAS example; this is acceptable, but the authors should ensure it is clearly positioned as an example, not as a substitute for a broader literature discussion.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: submitted body is a TinyML survey with external benchmarks; the abstract's SHeRL-FL results are unsupported but not circular.

full rationale

The submitted full text is not the paper announced in the abstract. The abstract claims SHeRL-FL cuts data transmission by over 90% vs centralized FL/HierFL and 50% vs SplitFed, with experiments on CIFAR-10/100, HAM10000, and ISIC-2018. The body is 'Designing Object Detection Models for TinyML' by different authors and contains no SHeRL-FL, split learning, hierarchical FL, or those experiments. That is a missing-evidence/completeness failure for the abstract's quantitative claim, but it is not circularity: there is no equation, fitted parameter, or self-citation by which the claimed reduction is defined into existence. The survey that is actually present is externally grounded: its comparative tables are compiled from cited papers (Table 5 states 'All results are compiled from the corresponding papers'), its taxonomy of quantization, pruning, knowledge distillation, and NAS is standard, and its equations are textbook definitions. The only self-citations, [35] and [164], are used as examples (e.g., the HNAS/LLM-NAS framework in Section 6.4.4) and are not load-bearing premises for any central claim. Hence the derivation chain, such as it is, contains no step that reduces to its own input.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The survey introduces no free parameters or invented entities. Its reliability rests on two domain assumptions: the chosen scope of optimization techniques is the right one, and the external numbers quoted in the tables are accurate transcriptions.

assumptions (2)
  • domain assumption The surveyed topics (quantization, pruning, KD, NAS) and the chosen hardware platforms constitute the comprehensive space for TinyML object detection.
    The survey's gap-filling claim in Section 1 depends on this coverage being the right scope; the paper does not justify completeness beyond prior survey comparison.
  • domain assumption Tabulated performance numbers (Tables 5, 6, 8) are accurately transcribed from cited works.
    The comparative analysis relies on external reported metrics, which the authors do not independently measure.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SHeRL-FL: When Representation Learning Meets Split Learning in Hierarchical Federated Learning." pith.science (2026). https://pith.science/paper/74P3OX7X

@misc{pith2026250808339,
  author       = {Pith},
  title        = {Pith review of: SHeRL-FL: When Representation Learning Meets Split Learning in Hierarchical Federated Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/74P3OX7X}},
  note         = {Machine review of arXiv:2508.08339}
}
read the original abstract

Federated learning (FL) is a promising approach for addressing scalability and latency issues in large-scale networks by enabling collaborative model training without requiring the sharing of raw data. However, existing FL frameworks often overlook the computational heterogeneity of edge clients and the growing training burden on resource-limited devices. However, FL suffers from high communication costs and complex model aggregation, especially with large models. Previous works combine split learning (SL) and hierarchical FL (HierFL) to reduce device-side computation and improve scalability, but this introduces training complexity due to coordination across tiers. To address these issues, we propose SHeRL-FL, which integrates SL and hierarchical model aggregation and incorporates representation learning at intermediate layers. By allowing clients and edge servers to compute training objectives independently of the cloud, SHeRL-FL significantly reduces both coordination complexity and communication overhead. To evaluate the effectiveness and efficiency of SHeRL-FL, we performed experiments on image classification tasks using CIFAR-10, CIFAR-100, and HAM10000 with AlexNet, ResNet-18, and ResNet-50 in both IID and non-IID settings. In addition, we evaluate performance on image segmentation tasks using the ISIC-2018 dataset with a ResNet-50-based U-Net. Experimental results demonstrate that SHeRL-FL reduces data transmission by over 90\% compared to centralized FL and HierFL, and by 50\% compared to SplitFed, which is a hybrid of FL and SL, and further improves hierarchical split learning methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. QSplitFL: Capability Aware Deep Q-Learning for Optimal Split Point Selection in Split Federated Learning

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    QSplitFL is a DQN framework that selects split points in split federated learning from hardware metrics with a decayed loss-drop reward and committee voting, reporting faster convergence and higher accuracy than basel...

Reference graph

Works this paper leans on

177 extracted references · 40 canonical work pages · cited by 1 Pith paper

  1. [1]

    Andrea Giovanni Accettola and Massimo Merenda. 2023. Dataset distillation as an enabling technique for on-device training in TinyML for IoT: an RFID use case. In ���� ��� ������������� ���������� �� ����� ��� ����������� ������������ ����������. 1–4

  2. [2]

    Shay Aharon, Louis-Dupont, Ofri Masad, Kate Yurkova, Lotem Fridman, Lkdci, Eugene Khvedchenya, Ran Rubin, Natan Bagrov, Borys Tymchenko, Tomer Keren, Alexander Zhilko, and Eran-Deci. 2021. Super-Gradients

  3. [3]

    Daghash K Alqahtani, Muhammad Aamir Cheema, and Adel N Toosi. 2024. Benchmarking deep learning models for object detection on edge computing devices. In ������������� ���������� �� ���������������� ���������. Springer, 142–150

  4. [4]

    Alberto Ancilotto, Francesco Paissan, and Elisabetta Farella. 2023. XiNet: Efficient Neural Networks for tinyML. In ����������� �� ��� �������� ������������� ���������� �� �������� ������ ������. 16968–16977

  5. [5]

    Colby Banbury, Emil Njor, Matthew Stewart, Pete Warden, Manjunath Kudlur, Nat Jeffries, Xenofon Fafoutis, and Vijay Janapa Reddi. 2024. Wake Vision: A Large-scale, Diverse Dataset and Benchmark Suite for TinyML Person Detection. ����� ��������, Article arXiv:2405.00892 (May 2024), arXiv:2405.00892 pages. arXiv:2405.00892 [cs.CV]

  6. [6]

    Colby Banbury, Vijay Janapa Reddi, Peter Torelli, Jeremy Holleman, Nat Jeffries, Csaba Kiraly, Pietro Montino, David Kanter, Sebastian Ahmed, Danilo Pau, et al. 2021. Mlperf tiny benchmark. ����� �������� ����������������(2021)

  7. [7]

    Banbury, Chuteng Zhou, Igor Fedorov, Ramon Matas Navarro, Urmish Thakker, Dibakar Gope, Vijay Janapa Reddi, Matthew Mattina, and Paul N

    Colby R. Banbury, Chuteng Zhou, Igor Fedorov, Ramon Matas Navarro, Urmish Thakker, Dibakar Gope, Vijay Janapa Reddi, Matthew Mattina, and Paul N. Whatmough. 2020. MicroNets: Neural Network Architectures for Deploying TinyML Applications on Commodity Microcontrollers. ���� abs/2010.11267 (2020). arXiv:2010.11267

  8. [8]

    Amin Banitalebi-Dehkordi. 2021. Knowledge Distillation for Low-Power Object Detection: A Simple Technique and Its Extensions for Training Compact Models Using Unlabeled Data. In ���� �������� ������������� ���������� �� �������� ������ ��������� �������. 769–778

Show all 177 references
  1. [9]

    Amin Banitalebi-Dehkordi. 2021. Revisiting Knowledge Distillation for Object Detection. ���� abs/2105.10633 (2021). arXiv:2105.10633

  2. [10]

    Gabriel Bender, Hanxiao Liu, Bo Chen, Grace Chu, Shuyang Cheng, Pieter-Jan Kindermans, and Quoc Le. 2020. Can weight sharing outperform random architecture search? An investigation with TuNAS. ���� abs/2008.06120 (2020). arXiv:2008.06120

  3. [11]

    Courville

    Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville. 2013. Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation. ���� abs/1308.3432 (2013). arXiv:1308.3432 Manuscript submitted to ACM 40 El Zeinaty et al

  4. [12]

    Pietro Bonazzi, Thomas Rüegg, Sizhen Bian, Yawei Li, and Michele Magno. 2023. TinyTracker: Ultra-Fast and Ultra-Low-Power Edge Vision In-Sensor for Gaze Estimation. In ���� ���� �������. 1–4

  5. [13]

    Han Cai, Chuang Gan, and Song Han. 2019. Once for All: Train One Network and Specialize it for Efficient Deployment. ���� abs/1908.09791 (2019). arXiv:1908.09791

  6. [14]

    Han Cai, Ligeng Zhu, and Song Han. 2018. ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware. ���� abs/1812.00332 (2018). arXiv:1812.00332

  7. [15]

    Yuxuan Cai, Hongjia Li, Geng Yuan, Wei Niu, Yanyu Li, Xulong Tang, Bin Ren, and Yanzhi Wang. 2020. YOLObile: Real-Time Object Detection on Mobile Devices via Compression-Compilation Co-Design. ����� abs/2009.05697 (2020)

  8. [16]

    Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-End Object Detection with Transformers. ���� abs/2005.12872 (2020). arXiv:2005.12872

  9. [17]

    Bo Chen, Golnaz Ghiasi, Hanxiao Liu, Tsung-Yi Lin, Dmitry Kalenichenko, Hartwig Adam, and Quoc V. Le. 2019. MnasFPN: Learning Latency-aware Pyramid Architecture for Object Detection on Mobile Devices. ���� abs/1912.01106 (2019). arXiv:1912.01106

  10. [18]

    Guobin Chen, Wongun Choi, Xiang Yu, Tony Han, and Manmohan Chandraker. 2017. Learning Efficient Object Detection Models with Knowledge Distillation. In �������� �� ������ ����������� ���������� �������, I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanatha...

  11. [19]

    Doermann, and Guodong Guo

    Hanlin Chen, Li’an Zhuo, Baochang Zhang, Xiawu Zheng, Jianzhuang Liu, Rongrong Ji, David S. Doermann, and Guodong Guo. 2020. Binarized Neural Architecture Search for Efficient Object Recognition. ���� abs/2009.04247 (2020). arXiv:2009.04247

  12. [20]

    Shaoyu Chen, Tianheng Cheng, Jiemin Fang, Qian Zhang, Yuan Li, Wenyu Liu, and Xinggang Wang. 2023. TinyDet: Accurate Small Object Detection in Lightweight Generic Detectors. arXiv:2304.03428 [cs.CV]

  13. [21]

    Shoufa Chen, Peize Sun, Yibing Song, and Ping Luo. 2022. DiffusionDet: Diffusion Model for Object Detection.����� ��������, Article arXiv:2211.09788 (Nov. 2022), arXiv:2211.09788 pages. arXiv:2211.09788 [cs.CV]

  14. [22]

    Xingjian Chen, Jianbo Su, and Jun Zhang. 2019. A Two-Teacher Framework for Knowledge Distillation. In �������� �� ������ �������� � ���� ����, Huchuan Lu, Huajin Tang, and Zhanshan Wang (Eds.). Springer International Publishing, Cham, 58–66

  15. [23]

    Yukang Chen, Tong Yang, Xiangyu Zhang, Gaofeng Meng, Chunhong Pan, and Jian Sun. 2019. DetNAS: Neural Architecture Search on Object Detection. ���� abs/1903.10979 (2019). arXiv:1903.10979

  16. [24]

    Tianheng Cheng, Lin Song, Yixiao Ge, Wenyu Liu, Xinggang Wang, and Ying Shan. 2024. Yolo-world: Real-time open-vocabulary object detection. In ����������� �� ��� �������� ���������� �� �������� ������ ��� ������� �����������. 16901–16911

  17. [25]

    Colwell, and Adrian Weller

    Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamás Sarlós, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy J. Colwell, and Adrian Weller. 2020. Rethinking Attention with Performers. ���� abs/2009.14794 ...

  18. [26]

    Aakanksha Chowdhery, Pete Warden, Jonathon Shlens, Andrew Howard, and Rocky Rhodes. 2019. Visual Wake Words Dataset.���� abs/1906.05721 (2019). arXiv:1906.05721

  19. [27]

    Viviana Crescitelli, Seiji Miura, Goichi Ono, and Naohiro Kohmu. 2021. Edge devices object detection by filter pruning. In���� ���� ���� ������������� ���������� �� �������� ������������ ��� ������� ���������� ����� �. 1–7

  20. [28]

    Li Cuimei, Qi Zhiliang, Jia Nan, and Wu Jianhua. 2017. Human face detection algorithm via Haar cascade classifier combined with three additional classifiers. In ���� ���� ���� ������������� ���������� �� ���������� ����������� � ����������� �������. 483–487

  21. [29]

    Dalal and B

    N. Dalal and B. Triggs. 2005. Histograms of oriented gradients for human detection. In ���� ���� �������� ������� ���������� �� �������� ������ ��� ������� ����������� ���������, Vol. 1. 886–893 vol. 1

  22. [30]

    Josen Daniel De Leon and Rowel Atienza. 2022. Depth pruning with auxiliary networks for tinyml. In ������ ��������� ���� ������������� ���������� �� ���������� ������ ��� ������ ���������� ��������. IEEE, 3963–3967

  23. [31]

    Jieren Deng, Xin Zhou, Hao Tian, Zhihong Pan, and Derek Aguiar. 2023. Smooth and Stepwise Self-Distillation for Object Detection. ����� ��������, Article arXiv:2303.05015 (March 2023), arXiv:2303.05015 pages. arXiv:2303.05015 [cs.CV]

  24. [32]

    Caiwen Ding, Shuo Wang, Ning Liu, Kaidi Xu, Yanzhi Wang, and Yun Liang. 2019. REQ-YOLO: A Resource-Aware, Efficient Quantization Framework for Object Detection on FPGAs. ���� abs/1909.13396 (2019). arXiv:1909.13396

  25. [33]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognit...

  26. [34]

    Kaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi, Qingming Huang, and Qi Tian. 2019. CenterNet: Keypoint Triplets for Object Detection. ���� abs/1904.08189 (2019). arXiv:1904.08189

  27. [35]

    Christophe El Zeinaty, Glenn Herrou, Wassim Hamidouche, and Daniel Menard. 2024. Dicetrack: Lightweight Dice Classification on Resource- Constrained Platforms with Optimized Deep Learning Models. In ������ ���� � ���� ���� ������������� ���������� �� ���������� ������ ��� ����...

  28. [36]

    Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. 2019. Neural Architecture Search: A Survey.�� ����� ������ ����20, 1 (jan 2019), 1997–2017

  29. [37]

    Erik Englesson and Hossein Azizpour. 2021. Generalized Jensen-Shannon Divergence Loss for Learning with Noisy Labels. ���� abs/2105.04522 (2021). arXiv:2105.04522 Manuscript submitted to ACM Designing Object Detection Models for TinyML: Foundations, Comparative Analysis, Chall...

  30. [38]

    Everingham, L

    M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. [n. d.]. The PASCAL Visual Object Classes Challenge 2007 (VOC2007) Results. http://www.pascal-network.org/challenges/VOC/voc2007/workshop/index.html

  31. [39]

    Jiemin Fang, Yuzhu Sun, Qian Zhang, Yuan Li, Wenyu Liu, and Xinggang Wang. 2019. Densely Connected Search Space for More Flexible Neural Architecture Search. ���� abs/1906.09607 (2019). arXiv:1906.09607

  32. [40]

    Jiemin Fang, Yuzhu Sun, Qian Zhang, Kangjian Peng, Yuan Li, Wenyu Liu, and Xinggang Wang. 2020. FNA++: Fast Network Adaptation via Parameter Remapping and Architecture Search. ���� abs/2006.12986 (2020). arXiv:2006.12986

  33. [41]

    Felzenszwalb, Ross B

    Pedro F. Felzenszwalb, Ross B. Girshick, David McAllester, and Deva Ramanan. 2010. Object Detection with Discriminatively Trained Part-Based Models. ���� ������������ �� ������� �������� ��� ������� ������������32, 9 (2010), 1627–1645

  34. [42]

    Xin Feng, Youni Jiang, Xuejiao Yang, Ming Du, and Xin Li. 2019. Computer vision algorithms and hardware implementations: A survey. ����������� 69 (2019), 309–320

  35. [43]

    Jiyang Gao, Jiang Wang, Shengyang Dai, Li-Jia Li, and Ram Nevatia. 2018. NOTE-RCNN: NOise Tolerant Ensemble RCNN for Semi-Supervised Object Detection. ���� abs/1812.00124 (2018). arXiv:1812.00124

  36. [44]

    Tao Ge, Si-Qing Chen, and Furu Wei. 2022. EdgeFormer: A parameter-efficient transformer for on-device Seq2Seq generation. ����� �������� ����������������(2022)

  37. [45]

    Golnaz Ghiasi, Tsung-Yi Lin, Ruoming Pang, and Quoc V. Le. 2019. NAS-FPN: Learning Scalable Feature Pyramid Architecture for Object Detection. ���� abs/1904.07392 (2019). arXiv:1904.07392

  38. [46]

    Lee Giles, Gary M

    C. Lee Giles, Gary M. Kuhn, and Ronald J. Williams. 1994. Dynamic recurrent neural networks: Theory and applications. ���� ������������ �� ������ ��������5, 2 (1994), 153–156

  39. [47]

    Ross Girshick, Pedro Felzenszwalb, and David McAllester. 2011. Object Detection with Grammar Models. In �������� �� ������ ����������� ���������� �������, J. Shawe-Taylor, R. Zemel, P. Bartlett, F. Pereira, and K.Q. Weinberger (Eds.), Vol. 24. Curran Associates, Inc

  40. [48]

    Cunhan Guo and Heyan Huang. 2025. Enhancing camouflaged object detection through contrastive learning and data augmentation techniques. ����������� ������������ �� ��������� ������������141 (2025), 109703

  41. [49]

    Jianyuan Guo, Kai Han, Yunhe Wang, Chao Zhang, Zhaohui Yang, Han Wu, Xinghao Chen, and Chang Xu. 2020. Hit-Detector: Hierarchical Trinity Architecture Search for Object Detection. ���� abs/2003.11818 (2020). arXiv:2003.11818

  42. [50]

    Zhen Guo, Pengzhou Zhang, and Peng Liang. 2024. Shared Knowledge Distillation Network for Object Detection. �����������13, 8 (2024)

  43. [51]

    Mohammad Hajizadeh, Mohammad Sabokrou, and Adel Rahmani. 2023. MobileDenseNet: A new approach to object detection on mobile devices. ������ ������� ���� ������������215 (2023), 119348

  44. [52]

    Lixiang Han, Zhen Xiao, and Zhenjiang Li. 2024. Dtmm: Deploying tinyml models on extremely weak iot devices with pruning. In ���� ������� ��������� ���������� �� �������� ��������������. IEEE, 1999–2008

  45. [53]

    Jianwei Hao, Piyush Subedi, Lakshmish Ramaswamy, and In Kee Kim. 2023. Reaching for the Sky: Maximizing Deep Learning Inference Throughput on Edge Devices with AI Multi-Tenancy. ��� ������������ �� �������� ����������23, 1, Article 2 (Feb. 2023), 33 pages

  46. [54]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. Deep Residual Learning for Image Recognition. ���� abs/1512.03385 (2015). arXiv:1512.03385

  47. [55]

    Ryota Hinami and Shin’ichi Satoh. 2016. Large-Scale R-CNN with Classifier Adaptive Quantization. In�������� ������ � ���� ���� � ���� �������� ����������� ���������� ��� ������������ ������� ������ ����� ������������ ���� ��� �������� ����� �� �������� �������� ���� �����, Bas...

  48. [56]

    Geoffrey Hinton, Jeff Dean, and Oriol Vinyals. 2014. Distilling the Knowledge in a Neural Network. 1–9

  49. [57]

    Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam

    Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. ���� abs/1704.04861 (2017). arXiv:1704.04861

  50. [58]

    Zeyi Huang, Yang Zou, Vijayakumar Bhagavatula, and Dong Huang. 2020. Comprehensive Attention Self-Distillation for Weakly-Supervised Object Detection. ���� abs/2010.12023 (2020). arXiv:2010.12023

  51. [59]

    Edge Impulse. 2022. FOMO: Object Detection for Constrained Devices. Accessed: 02-01-2024

  52. [60]

    Howard, Hartwig Adam, and Dmitry Kalenichenko

    Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew G. Howard, Hartwig Adam, and Dmitry Kalenichenko. 2017. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. ���� abs/1712.05877 (2017). arXiv:1712.05877

  53. [61]

    Eunjin Jeong, Jangryul Kim, and Soonhoi Ha. 2022. TensorRT-Based Framework and Optimization Methodology for Deep Learning Inference on Jetson Boards. ��� ������ ������ ������� �����21, 5, Article 51 (oct 2022), 26 pages

  54. [62]

    Learned-Miller

    SouYoung Jin, Aruni RoyChowdhury, Huaizu Jiang, Ashish Singh, Aditya Prasad, Deep Chakraborty, and Erik G. Learned-Miller. 2018. Unsupervised Hard Example Mining from Videos for Improved Object Detection. ���� abs/1808.04285 (2018). arXiv:1808.04285

  55. [63]

    Glenn Jocher et al. 2022. ������������������� ���� � ������ ���� �������� �������� ������������

  56. [64]

    Glenn Jocher and Jing Qiu. 2024. ����������� ������

  57. [65]

    Do-Yoon Jung, Yeon-Jae Oh, and Nam-Ho Kim. 2024. A study on GAN-Based Car body part defect detection process and Comparative Analysis of YOLO v7 and YOLO v8 object detection performance. �����������13, 13 (2024), 2598

  58. [66]

    Pilsung Kang and Athip Somtham. 2022. An Evaluation of Modern Accelerator-Based Edge Devices for Object Detection Applications.����������� 10, 22 (2022)

  59. [67]

    Mandeep Kaur and Rajni Aron. 2021. A systematic study of load balancing approaches in the fog computing environment. ��� ������� �� ��������������77, 8 (Feb. 2021), 9202–9247. Manuscript submitted to ACM 42 El Zeinaty et al

  60. [68]

    Kyungho Kim, Sung-Joon Jang, Jonghee Park, Eunchong Lee, and Sang-Seol Lee. 2023. Lightweight and Energy-Efficient Deep Learning Accelerator for Real-Time Object Detection on Edge Devices. �������23, 3 (2023)

  61. [69]

    Seijoon Kim, Seongsik Park, Byunggook Na, and Sungroh Yoon. 2019. Spiking-YOLO: Spiking Neural Network for Energy-Efficient Object Detection. ����� ��������, Article arXiv:1903.06530 (March 2019), arXiv:1903.06530 pages. arXiv:1903.06530 [cs.CV]

  62. [70]

    Yuma Koizumi, Shoichiro Saito, Hisashi Uematsu, Noboru Harada, and Keisuke Imoto. 2019. ToyADMOS: A Dataset of Miniature-Machine Operating Sounds for Anomalous Sound Detection. ����� ��������, Article arXiv:1908.03299 (Aug. 2019), arXiv:1908.03299 pages. arXiv:1908.03299 [eess.AS]

  63. [71]

    Srinivas S. S. Kruthiventi, Pratyush Sahay, and Rajesh Biswal. 2017. Low-light pedestrian detection from RGB images using multi-modal knowledge distillation. In ���� ���� ������������� ���������� �� ����� ���������� ������. 4207–4211

  64. [72]

    Hongbo Kuang and Ziwei Liu. 2021. Research on Object Detection Network Based on Knowledge Distillation. In ���� ��� ������������� ���������� �� ����������� ���������� ������� ��������. 8–12

  65. [73]

    Jaeha Kung, David Zhang, Gooitzen Wal, Sek Chai, and Saibal Mukhopadhyay. 2018. Efficient Object Detection Using Embedded Binarized Neural Networks. �� ������ �������� �����90, 6 (jun 2018), 877–890

  66. [74]

    Mark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev, John Carr, Michael Goin, William Leiserson, Sage Moore, Nir Shavit, and Dan Alistarh. 2020. Inducing and Exploiting Activation Sparsity for Fast Inference on Deep Neural Networks. In ����������� �� ��� ���� �����...

  67. [75]

    Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper R. R. Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Tom Duerig, and Vittorio Ferrari. 2018. The Open Images Dataset V4: Unified image classification, object detection, and visual relatio...

  68. [76]

    Hei Law and Jia Deng. 2018. CornerNet: Detecting Objects as Paired Keypoints. ���� abs/1808.01244 (2018). arXiv:1808.01244

  69. [77]

    Yann LeCun, John Denker, and Sara Solla. 1989. Optimal Brain Damage. In �������� �� ������ ����������� ���������� �������, D. Touretzky (Ed.), Vol. 2. Morgan-Kaufmann

  70. [78]

    Desai, and Mooi Choo Chuah

    Dawei Li, Theodoros Salonidis, Nirmit V. Desai, and Mooi Choo Chuah. 2016. DeepCham: Collaborative Edge-Mediated Adaptive Deep Learning for Mobile Object Recognition. In ���� �������� ��������� �� ���� ��������� �����. 64–76

  71. [79]

    Lei Li, Alexander Linger, Mario Millhäusler, Vagia Tsiminaki, Yuanyou Li, and Dengxin Dai. 2024. Object-centric Cross-modal Feature Distillation for Event-based Object Detection. In ���� ���� ������������� ���������� �� �������� ��� ���������� ������. 15440–15447

  72. [80]

    Quanquan Li, Shengying Jin, and Junjie Yan. 2017. Mimicking Very Efficient Network for Object Detection. In ����������� �� ��� ���� ���������� �� �������� ������ ��� ������� ����������� ������

  73. [81]

    Rundong Li, Yan Wang, Feng Liang, Hongwei Qin, Junjie Yan, and Rui Fan. 2019. Fully Quantized Network for Object Detection. In ���� �������� ���������� �� �������� ������ ��� ������� ����������� ������. 2805–2814

  74. [82]

    Yuxi Li, Jiuwei Li, Weiyao Lin, and Jianguo Li. 2018. Tiny-DSOD: Lightweight Object Detection for Resource-Restricted Usages.���� abs/1807.11013 (2018). arXiv:1807.11013

  75. [83]

    Ming Liang and Xiaolin Hu. 2015. Recurrent convolutional neural network for object recognition. In ���� ���� ���������� �� �������� ������ ��� ������� ����������� ������. 3367–3375

  76. [84]

    Siyuan Liang, Hao Wu, Li Zhen, Qiaozhi Hua, Sahil Garg, Georges Kaddoum, Mohammad Mehedi Hassan, and Keping Yu. 2022. Edge YOLO: Real-Time Intelligent Object Detection System Based on Edge-Cloud Cooperation in Autonomous Vehicles. ���� ������������ �� ����������� �������������...

  77. [85]

    Yinan Liang, Ziwei Wang, Xiuwei Xu, Yansong Tang, Jie Zhou, and Jiwen Lu. 2023. Mcuformer: Deploying vision tranformers on microcontrollers with limited memory. �������� �� ������ ����������� ���������� �������36 (2023), 8501–8512

  78. [86]

    Benedetta Liberatori, Ciro Antonio Mami, Giovanni Santacatterina, Marco Zullich, and Felice Andrea Pellegrino. 2022. YOLO-Based Face Mask Detection on Low-End Devices Using Pruning and Quantization. In ���� ���� ������� ������������� ���������� �� ������������ ������������� ��...

  79. [87]

    Ji Lin, Wei-Ming Chen, Han Cai, Chuang Gan, and Song Han. 2021. MCUNetV2: memory-efficient patch-based inference for tiny deep learning. In ����������� �� ��� ���� ������������� ���������� �� ������ ����������� ���������� ������� ����� ����. Curran Associates Inc., Red Hook, N...

  80. [88]

    Ji Lin, Wei-Ming Chen, Yujun Lin, John Cohn, Chuang Gan, and Song Han. 2020. MCUNet: tiny deep learning on IoT devices. In ����������� �� ��� ���� ������������� ���������� �� ������ ����������� ���������� �������(Vancouver, BC, Canada) ����� ����. Curran Associates Inc., Red H...

  81. [89]

    Ji Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang, and Song Han. 2024. Tiny Machine Learning: Progress and Futures. ����� ��������, Article arXiv:2403.19076 (March 2024), arXiv:2403.19076 pages. arXiv:2403.19076 [cs.LG]

  82. [90]

    Girshick, Kaiming He, Bharath Hariharan, and Serge J

    Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie. 2016. Feature Pyramid Networks for Object Detection. ���� abs/1612.03144 (2016). arXiv:1612.03144

  83. [91]

    Girshick, Kaiming He, and Piotr Dollár

    Tsung-Yi Lin, Priya Goyal, Ross B. Girshick, Kaiming He, and Piotr Dollár. 2017. Focal Loss for Dense Object Detection. ���� abs/1708.02002 (2017). arXiv:1708.02002

  84. [92]

    Belongie, Lubomir D

    Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context. ���� abs/1405.0312 (2014). arXiv:1405.0312

  85. [93]

    Hanxiao Liu, Karen Simonyan, and Yiming Yang. 2018. DARTS: Differentiable Architecture Search. ���� abs/1806.09055 (2018). arXiv:1806.09055 Manuscript submitted to ACM Designing Object Detection Models for TinyML: Foundations, Comparative Analysis, Challenges, and Emerging Sol...

  86. [94]

    Hou-I Liu, Marco Galindo, Hongxia Xie, Lai-Kuan Wong, Hong-Han Shuai, Yung-Hui Li, and Wen-Huang Cheng. 2024. Lightweight Deep Learning for Resource-Constrained Environments: A Survey. ����� ��������, Article arXiv:2404.07236 (April 2024), arXiv:2404.07236 pages. arXiv:2404.07...

  87. [95]

    Jie Liu, Chuming Li, Feng Liang, Chen Lin, Ming Sun, Junjie Yan, Wanli Ouyang, and Dong Xu. 2020. Inception Convolution with Efficient Dilation Search. ���� abs/2012.13587 (2020). arXiv:2012.13587

  88. [96]

    Kai Liu, Zhihang Fu, Sheng Jin, Ze Chen, Fan Zhou, Rongxin Jiang, Yaowu Chen, and Jieping Ye. 2024. ESOD: Efficient Small Object Detection on High-Resolution Images. ���� ������������ �� ����� ����������(2024)

  89. [97]

    Fieguth, Jie Chen, Xinwang Liu, and Matti Pietikäinen

    Li Liu, Wanli Ouyang, Xiaogang Wang, Paul W. Fieguth, Jie Chen, Xinwang Liu, and Matti Pietikäinen. 2018. Deep Learning for Generic Object Detection: A Survey. ���� abs/1809.02165 (2018). arXiv:1809.02165

  90. [98]

    Shu Liu, Lu Qi, Haifang Qin, Jianping Shi, and Jiaya Jia. 2018. Path Aggregation Network for Instance Segmentation. ���� abs/1803.01534 (2018). arXiv:1803.01534

  91. [99]

    Reed, Cheng-Yang Fu, and Alexander C

    Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott E. Reed, Cheng-Yang Fu, and Alexander C. Berg. 2015. SSD: Single Shot MultiBox Detector. ���� abs/1512.02325 (2015). arXiv:1512.02325

  92. [100]

    Yuanpei Liu, Xingping Dong, Wenguan Wang, and Jianbing Shen. 2019. Teacher-Students Knowledge Distillation for Siamese Trackers. ����� abs/1907.10586 (2019)

  93. [101]

    Zhuoming Liu, Xuefeng Hu, and Ram Nevatia. 2024. Efficient Feature Distillation for Zero-Shot Annotation Object Detection. In ����������� �� ��� �������� ������ ���������� �� ������������ �� �������� ������ ������. 893–902

  94. [102]

    Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. 2017. Learning Efficient Convolutional Networks Through Network Slimming. In ����������� �� ��� ���� ������������� ���������� �� �������� ������ ������

  95. [103]

    Zechun Liu, Baoyuan Wu, Wenhan Luo, Xin Yang, Wei Liu, and Kwang-Ting Cheng. 2018. Bi-Real Net: Enhancing the Performance of 1-bit CNNs With Improved Representational Capability and Advanced Training Algorithm. ���� abs/1808.00278 (2018). arXiv:1808.00278

  96. [104]

    Siliang Ma and Yong Xu. 2023. MPDIoU: A Loss for Efficient and Accurate Bounding Box Regression. ����� ��������, Article arXiv:2307.07662 (July 2023), arXiv:2307.07662 pages. arXiv:2307.07662 [cs.CV]

  97. [105]

    Sachin Mehta and Mohammad Rastegari. 2021. MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer. ���� abs/2110.02178 (2021). arXiv:2110.02178

  98. [106]

    T. Mita, T. Kaneko, and O. Hori. 2005. Joint Haar-like features for face detection. In����� ���� ������������� ���������� �� �������� ������ ��������� ������ �, Vol. 2. 1619–1626 Vol. 2

  99. [107]

    Payal Mittal. 2024. A comprehensive survey of deep learning-based lightweight object detection models for edge devices. ��������� ������������ ������57 (08 2024)

  100. [108]

    Pavlo Molchanov, Arun Mallya, Stephen Tyree, Iuri Frosio, and Jan Kautz. 2019. Importance Estimation for Neural Network Pruning. In ���� �������� ���������� �� �������� ������ ��� ������� ����������� ������. 11256–11264

  101. [109]

    Julian Moosmann, Pietro Bonazzi, Yawei Li, Sizhen Bian, Philipp Mayer, Luca Benini, and Michele Magno. 2023. Ultra-efficient on-device object detection on ai-integrated smart glasses with tinyissimoyolo. ����� �������� ����������������(2023)

  102. [110]

    Julian Moosmann, Hanna Müller, Nicky Zimmerman, Georg Rutishauser, Luca Benini, and Michele Magno. 2024. Flexible and Fully Quantized Lightweight TinyissimoYOLO for Ultra-Low-Power Edge Systems. ���� ������12 (2024), 75093–75107

  103. [111]

    Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart van Baalen, and Tijmen Blankevoort. 2021. A White Paper on Neural Network Quantization. ���� abs/2106.08295 (2021). arXiv:2106.08295

  104. [112]

    Quang-Huy Nguyen, Jin Peng Zhou, Zhenzhen Liu, Khanh-Huyen Bui, Kilian Q Weinberger, and Dung D Le. 2024. Zero-shot Object-Level OOD Detection with Context-Aware Inpainting. ����� �������� ����������������(2024)

  105. [113]

    Emre Ozer, Jedrzej Kufel, Shvetank Prakash, Alireza Raisiardali, Olof Kindgren, Ronald Wong, Nelson Ng, Damien Jausseran, Feras Alkhalil, David Kong, Gage Hills, Richard Price, and Vijay Janapa Reddi. 2024. Bendable non-silicon RISC-V microprocessor. ������(25 Sep 2024)

  106. [114]

    Francesco Paissan, Alberto Ancilotto, and Elisabetta Farella. 2021. PhiNets: a scalable backbone for low-power AI at the edge. ���� abs/2110.00337 (2021). arXiv:2110.00337

  107. [115]

    Patterson and John L

    David A. Patterson and John L. Hennessy. 2017. �������� ������������ ��� ������ ������ �������� ��� �������� �������� ���������(1st ed.). Morgan Kaufmann Publishers Inc., San Francisco, CA, USA

  108. [116]

    Hanyu Peng and Shifeng Chen. 2019. BDNN: Binary convolution neural networks for fast object detection. ������� ����������� �������125 (2019), 91–97

  109. [117]

    Junran Peng, Ming Sun, Zhaoxiang Zhang, Tieniu Tan, and Junjie Yan. 2019. Efficient Neural Architecture Transformation Searchin Channel-Level for Object Detection. ���� abs/1909.02293 (2019). arXiv:1909.02293

  110. [118]

    Hoang-The Pham, Minh-Anh Nguyen, and Chi-Chia Sun. 2019. AIoT Solution Survey and Comparison in Machine Learning on Low-cost Microcontroller. In ���� ������������� ��������� �� ����������� ������ ���������� ��� ������������� ������� ��������. 1–2

  111. [119]

    Plana, David Clark, Simon Davidson, Steve Furber, Jim Garside, Eustace Painkras, Jeffrey Pepper, Steve Temple, and John Bainbridge

    Luis A. Plana, David Clark, Simon Davidson, Steve Furber, Jim Garside, Eustace Painkras, Jeffrey Pepper, Steve Temple, and John Bainbridge. 2011. SpiNNaker: Design and Implementation of a GALS Multicore System-on-Chip. �� ������ �������� ������� �����7, 4, Article 17 (dec 2011...

  112. [120]

    Green, Pete Warden, Tim Ansell, and Vijay Janapa Reddi

    Shvetank Prakash, Tim Callahan, Joseph Bushagour, Colby Banbury, Alan V. Green, Pete Warden, Tim Ansell, and Vijay Janapa Reddi. 2023. CFU Playground: Full-Stack Open-Source Framework for Tiny Machine Learning (TinyML) Acceleration on FPGAs. In ���� ���� ������������� ��������...

  113. [121]

    Jinye Qu, Zeyu Gao, Tielin Zhang, Yanfeng Lu, Huajin Tang, and Hong Qiao. 2023. Spiking Neural Network for Ultra-low-latency and High-accurate Object Detection. arXiv:2306.12010 [cs.CV]

  114. [122]

    Rachmanto, Zaki Sukma, Ahmad N

    Rakandhiya D. Rachmanto, Zaki Sukma, Ahmad N. L. Nabhaan, Arief Setyanto, Ting Jiang, and In Kee Kim. 2024. Characterizing Deep Learning Model Compression with Post-Training Quantization on Accelerated Edge Devices. In ���� ���� ������������� ���������� �� ���� ��������� ��� �...

  115. [123]

    Girshick, and Ali Farhadi

    Joseph Redmon, Santosh Kumar Divvala, Ross B. Girshick, and Ali Farhadi. 2015. You Only Look Once: Unified, Real-Time Object Detection. ���� abs/1506.02640 (2015). arXiv:1506.02640

  116. [124]

    Girshick, and Jian Sun

    Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. ���� abs/1506.01497 (2015). arXiv:1506.01497

  117. [125]

    Tianhe Ren, Qing Jiang, Shilong Liu, Zhaoyang Zeng, Wenlong Liu, Han Gao, Hongjie Huang, Zhengyu Ma, Xiaoke Jiang, Yihao Chen, et al. 2024. Grounding dino 1.5: Advance the" edge" of open-set object detection. ����� �������� ����������������(2024)

  118. [126]

    Francisco Rivera Valverde, Juana Valeria Hurtado, and Abhinav Valada. 2021. There is More than Meets the Eye: Self-Supervised Multi-Object Detec- tion and Tracking with Sound by Distilling Multimodal Knowledge. ����� ��������, Article arXiv:2103.01353 (March 2021), arXiv:2103....

  119. [127]

    Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2015. FitNets: Hints for Thin Deep Nets. In ��� ������������� ���������� �� �������� ���������������� ���� ����� ��� ������ ��� ���� ��� ���� ����� ���������� ����� �������...

  120. [128]

    Bernstein, Alexander C

    Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Li Fei-Fei. 2014. ImageNet Large Scale Visual Recognition Challenge. ���� abs/1409.0575 (2014). arXiv:1409.0575

  121. [130]

    Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen

    Mark Sandler, Andrew G. Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. 2018. Inverted Residuals and Linear Bottlenecks: Mobile Networks for Classification, Detection and Segmentation. ���� abs/1801.04381 (2018). arXiv:1801.04381

  122. [131]

    Arief Setyanto, Theopilus Bayu Sasongko, Muhammad Ainul Fikri, and In Kee Kim. 2024. Near-Edge Computing Aware Object Detection: A Review. ���� ������12 (2024), 2989–3011

  123. [132]

    Nilotpal Sinha, Peyman Rostami, Abd El Rahman Shabayek, Anis Kacem, and Djamila Aouada. 2024. Multi-Objective Hardware Aware Neural Archi- tecture Search using Hardware Cost Diversity.����� ��������, Article arXiv:2404.12403 (April 2024), arXiv:2404.12403 pages. arXiv:2404.124...

  124. [133]

    Yang Song and Stefano Ermon. 2019. Generative Modeling by Estimating Gradients of the Data Distribution. In �������� �� ������ ����������� ���������� �������, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Eds.), Vol. 32. Curran Associates, Inc

  125. [134]

    Kai Su, Chowdhury MD Intisar, Qiangfu Zhao, and Yoichi Tomioka. 2020. Knowledge Distillation for Real-time On-Road Risk Detection. In ���� ���� ���� ���� �� ����������� ��������� ��� ������ ���������� ���� ���� �� ��������� ������������ ��� ���������� ���� ���� �� ����� ��� ��...

  126. [135]

    Baohua Sun, Tao Zhang, Jiapeng Su, and Hao Sha. 2021. GnetDet: Object Detection Optimized on a 224mW CNN Accelerator Chip at the Speed of 106FPS. ���� abs/2103.15756 (2021). arXiv:2103.15756

  127. [136]

    Siyang Sun, Yingjie Yin, Xingang Wang, De Xu, Wenqi Wu, and Qingyi Gu. 2018. Fast object detection based on binary deep convolution neural networks. ���� ������������ �� ������������ ����������3, 4 (2018), 191–197

  128. [137]

    Febin Sunny, Mahdi Nikdast, and Sudeep Pasricha. 2022. SONIC: A Sparse Neural Network Inference Accelerator with Silicon Photonics for Energy-Efficient Deep Learning. In ���� ���� ���� ��� ����� ������ ������ ���������� ���������� ���������. 214–219

  129. [138]

    Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour. 1999. Policy Gradient Methods for Reinforcement Learning with Function Approximation. In �������� �� ������ ����������� ���������� �������, S. Solla, T. Leen, and K. Müller (Eds.), Vol. 12. MIT Press

  130. [139]

    Vivienne Sze, Yu-Hsin Chen, Tien-Ju Yang, and Joel S. Emer. 2020. How to Evaluate Deep Neural Network Processors: TOPS/W (Alone) Considered Harmful. ���� ����������� �������� ��������12, 3 (2020), 28–41

  131. [140]

    Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, and Quoc V. Le. 2018. MnasNet: Platform-Aware Neural Architecture Search for Mobile. ���� abs/1807.11626 (2018). arXiv:1807.11626

  132. [141]

    Hidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, and Surya Ganguli. 2020. Pruning neural networks without any data by iteratively conserving synaptic flow. ���� abs/2006.05467 (2020). arXiv:2006.05467

  133. [142]

    Qiankun Tang, Jie Li, Zhiping Shi, and Yu Hu. 2020. Lightdet: A Lightweight and Accurate Object Detection Network. In ������ ���� � ���� ���� ������������� ���������� �� ���������� ������ ��� ������ ���������� ��������. 2243–2247

  134. [143]

    ManyCore Research Team. 2025. SpatialLM: Large Language Model for Spatial Understanding. https://github.com/manycore-research/SpatialLM

  135. [144]

    Juan Terven and Diana Cordova-Esparza. 2023. A Comprehensive Review of YOLO Architectures in Computer Vision: From YOLOv1 to YOLOv8 and YOLO-NAS. ����� ��������, Article arXiv:2304.00501 (April 2023), arXiv:2304.00501 pages. arXiv:2304.00501 [cs.CV]

  136. [145]

    Yunjie Tian, Qixiang Ye, and David Doermann. 2025. Yolov12: Attention-centric real-time object detectors. ����� �������� ����������������(2025)

  137. [146]

    Antonio Torralba. 2003. Contextual Priming for Object Detection. ������������� ������� �� �������� ������53, 2 (2003), 169–191. Publisher: Kluwer Academic Publishers ISBN: 1573-1405

  138. [147]

    Viola and M

    P. Viola and M. Jones. 2001. Rapid object detection using a boosted cascade of simple features. In ����������� �� ��� ���� ���� �������� ������� ���������� �� �������� ������ ��� ������� ������������ ���� ����, Vol. 1. I–I. Manuscript submitted to ACM Designing Object Detectio...

  139. [148]

    McLachlan

    Cinzia Viroli and Geoffrey J. McLachlan. 2017. Deep Gaussian Mixture Models. ����� ��������, Article arXiv:1711.06929 (Nov. 2017), arXiv:1711.06929 pages. arXiv:1711.06929 [stat.ML]

  140. [149]

    Ao Wang, Lihao Liu, Hui Chen, Zijia Lin, Jungong Han, and Guiguang Ding. 2025. YOLOE: Real-Time Seeing Anything. ����� �������� ����������������(2025)

  141. [150]

    Chaoqi Wang, Guodong Zhang, and Roger Baker Grosse. 2020. Picking Winning Tickets Before Training by Preserving Gradient Flow. ���� abs/2002.07376 (2020). arXiv:2002.07376

  142. [151]

    Ching-Hao Wang, Kang-Yang Huang, Yi Yao, Jun-Cheng Chen, Hong-Han Shuai, and Wen-Huang Cheng. 2024. Lightweight Deep Learning: An Overview. ���� �������� ����������� ��������13, 4 (2024), 51–64

  143. [152]

    Jiabao Wang, Yuming Chen, Zhaohui Zheng, Xiang Li, Ming-Ming Cheng, and Qibin Hou. 2024. CrossKD: Cross-Head Knowledge Distillation for Object Detection. In ����������� �� ��� �������� ���������� �� �������� ������ ��� ������� ����������� ������. 16520–16530

  144. [153]

    Ning Wang, Yang Gao, Hao Chen, Peng Wang, Zhi Tian, and Chunhua Shen. 2019. NAS-FCOS: Fast Neural Architecture Search for Object Detection. ���� abs/1906.04423 (2019). arXiv:1906.04423

  145. [154]

    Li, Madian Khabsa, Han Fang, and Hao Ma

    Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. 2020. Linformer: Self-Attention with Linear Complexity. ���� abs/2006.04768 (2020). arXiv:2006.04768

  146. [155]

    Xiaoxing Wang, Jiale Lin, Junchi Yan, Juanping Zhao, and Xiaokang Yang. 2022. EAutoDet: Efficient Architecture Search for Object Detection. ����� ��������, Article arXiv:2203.10747 (March 2022), arXiv:2203.10747 pages. arXiv:2203.10747 [cs.CV]

  147. [156]

    Ziwei Wang, Ziyi Wu, Jiwen Lu, and Jie Zhou. 2020. BiDet: An Efficient Binarized Object Detector. In ����������� �� ��� �������� ���������� �� �������� ������ ��� ������� ����������� ������

  148. [157]

    Pete Warden. 2018. Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition. ���� abs/1804.03209 (2018). arXiv:1804.03209

  149. [158]

    Kan Wu, Jinnian Zhang, Houwen Peng, Mengchen Liu, Bin Xiao, Jianlong Fu, and Lu Yuan. 2022. Tinyvit: Fast pretraining distillation for small vision transformers. In �������� ���������� �� �������� ������. Springer, 68–85

  150. [159]

    Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He

    Saining Xie, Ross B. Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. 2016. Aggregated Residual Transformations for Deep Neural Networks. ���� abs/1611.05431 (2016). arXiv:1611.05431

  151. [160]

    Yunyang Xiong, Hanxiao Liu, Suyog Gupta, Berkin Akin, Gabriel Bender, Pieter-Jan Kindermans, Mingxing Tan, Vikas Singh, and Bo Chen. 2020. MobileDets: Searching for Object Detection Architectures for Mobile Accelerators. ���� abs/2004.14525 (2020). arXiv:2004.14525

  152. [161]

    Jiaolong Xu, Peng Wang, Heng Yang, and Antonio M. López. 2018. Training a Binary Weight Object Detector by Knowledge Transfer for Autonomous Driving. ���� abs/1804.06332 (2018). arXiv:1804.06332

  153. [162]

    Kunran Xu, Yishi Li, Huawei Zhang, Rui Lai, and Lin Gu. 2022. EtinyNet: Extremely Tiny Network for TinyML. ����������� �� ��� ���� ���������� �� ��������� ������������36, 4 (Jun. 2022), 4628–4636

  154. [163]

    Syed Sahil Abbas Zaidi, Mohammad Samar Ansari, Asra Aslam, Nadia Kanwal, Mamoona Naveed Asghar, and Brian Lee. 2021. A Survey of Modern Deep Learning based Object Detection Models. ���� abs/2104.11892 (2021). arXiv:2104.11892

  155. [164]

    Christophe El Zeinaty, Wassim Hamidouche, Glenn Herrou, Daniel Menard, and Merouane Debbah. 2025. Can LLMs Revolutionize the Design of Explainable and Efficient TinyML Models? arXiv:2504.09685 [cs.LG]

  156. [165]

    Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu. 2024. Vision-language models for vision tasks: A survey. ���� ������������ �� ������� �������� ��� ������� ������������(2024)

  157. [166]

    Jiabin Zhang, Hu Su, Wei Zou, Xinyi Gong, Zhengtao Zhang, and Fei Shen. 2021. CADN: A weakly supervised learning-based category-aware object detection network for surface defect detection. ������� �����������109 (2021), 107571

  158. [167]

    Peizhen Zhang, Zijian Kang, Tong Yang, Xiangyu Zhang, Nanning Zheng, and Jian Sun. 2021. LGD: Label-guided Self-distillation for Object Detection. ���� abs/2109.11496 (2021). arXiv:2109.11496

  159. [168]

    Xingzhou Zhang, Yifan Wang, and Weisong Shi. 2019. pCAMP: Performance Comparison of Machine Learning Packages on the Edges. ���� abs/1906.01878 (2019). arXiv:1906.01878

  160. [169]

    Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. 2017. ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices. ���� abs/1707.01083 (2017). arXiv:1707.01083

  161. [170]

    Junhe Zhao, Sheng Xu, Runqi Wang, Baochang Zhang, Guodong Guo, David Doermann, and Dianmin Sun. 2022. Data-adaptive binary neural networks for efficient object detection and recognition. ������� ����������� �������153 (2022), 239–245

  162. [171]

    Yian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei, Guanzhong Wang, Qingqing Dang, Yi Liu, and Jie Chen. 2023. DETRs Beat YOLOs on Real-time Object Detection. ����� ��������, Article arXiv:2304.08069 (April 2023), arXiv:2304.08069 pages. arXiv:2304.08069 [cs.CV]

  163. [172]

    Mingkai Zheng, Xiu Su, Shan You, Fei Wang, Chen Qian, Chang Xu, and Samuel Albanie. 2023. Can GPT-4 Perform Neural Architecture Search? ����� ��������, Article arXiv:2304.10970 (April 2023), arXiv:2304.10970 pages. arXiv:2304.10970 [cs.LG]

  164. [173]

    Chuteng Zhou, Fernando García-Redondo, Julian Büchel, Irem Boybat, Xavier Timoneda Comas, S. R. Nandakumar, Shidhartha Das, Abu Sebastian, Manuel Le Gallo, and Paul N. Whatmough. 2021. AnalogNets: ML-HW Co-Design of Noise-robust TinyML Models and Always-On Analog Compute-in-Me...

  165. [174]

    Yipeng Zhou, Huaming Qian, and Peng Ding. 2023. Lite-YOLOv3: a real-time object detector based on multi-scale slice depthwise convolution and lightweight attention mechanism. �� ���� ���� ����� ��������20, 6 (2023), 123

  166. [175]

    Jingyuan Zhu, Shiyu Li, Yuxuan Andy Liu, Jian Yuan, Ping Huang, Jiulong Shan, and Huimin Ma. 2024. Odgen: Domain-specific object detection data generation with diffusion models. �������� �� ������ ����������� ���������� �������37 (2024), 63599–63633. Manuscript submitted to AC...

  167. [176]

    Bohan Zhuang, Chunhua Shen, Mingkui Tan, Lingqiao Liu, and Ian D. Reid. 2019. Structured Binary Neural Networks for Image Recognition. ���� abs/1909.09934 (2019). arXiv:1909.09934

  168. [177]

    Zhengxia Zou, Keyan Chen, Zhenwei Shi, Yuhong Guo, and Jieping Ye. 2023. Object Detection in 20 Years: A Survey. ����� ����111, 3 (2023), 257–276

  169. [178]

    Özge Ünel, Burak O

    F. Özge Ünel, Burak O. Özkalayci, and Cevahir Çiğla. 2019. The Power of Tiling for Small Object Detection. In ���� �������� ���������� �� �������� ������ ��� ������� ����������� ��������� �������. 582–591. Manuscript submitted to ACM

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.