REVIEW 3 major objections 5 minor 1 cited by
Collaborative Learning of On-Device Small Model and Cloud-Based Large Model: Advances and Future Directions
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This survey argues that cloud-only large-model learning has hit latency, cost, personalization, and privacy limits, and that on-device small models collaborating with cloud large models is the emerging fix.
desk verdict A capable survey with a genuinely useful taxonomy, marred by a 1000x dataset error and sloppy cost arithmetic that are fixable in review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing organizing device is a layered framework -- hardware, system and engine, model and algorithm, application -- together with a taxonomy of collaboration algorithms based on what is exchanged between cloud and device: data-based collaboration (raw samples or queries, including data filtering and query routing), feature-based collaboration (intermediate or final model outputs, including model splitting, early exiting, distillation, and parallel decoding), and parameter-based collaboration (models or model updates, including federated learning, model ensemble, and offsite or proxy tuning). This taxonomy is what lets the survey map dozens of works onto a small number of design choices: where subtasks are allocated, when exchange happens, and what is transmitted. The framework's second function is to show that the four identified bottlenecks each correspond to a layer where a concrete problem and a line of advances already exist.
What would settle it
Check the primary source [12] for the Reddit dataset's user count; if it is about 56,587 rather than 56,587,343, the survey's Table 2 contains a thousandfold error, and the same check should be run on the other dataset rows and on the industry cost estimate cited as [75] (about 30 million yuan per day for 300 million users).
Extended reading notes
Core claim
The paper's central claim is that collaborative learning between an on-device small model and a cloud-based large model can escape four bottlenecks of cloud-centric learning: latency too high for real-time interaction, cost and load too high when millions of devices upload raw data, a single global model that cannot personalize to individual users, and privacy risk from centralizing sensitive data. In the proposed paradigm, the small model handles local real-time inference, adapts to the user's data, and sends only non-sensitive samples, features, or parameter updates upward, while the large model transfers knowledge down via distillation or compression and handles complex global reasoning. The survey claims this creates a virtuous co-evolution cycle and supports it with a four-layer framework review and with industrial deployments in recommender systems, livestreaming content understanding, and personal intelligent assistants. It also argues that existing benchmarks fall short because they lack user-level or device-level partitioning and metrics.
Load-bearing premise
The survey's whole edifice rests on faithfully summarizing more than 180 cited works and industry reports; one concrete failure is visible already, since Table 2 lists the Reddit dataset as partitioned across 56,587,343 users while the cited benchmark [12] reports about 56,587 users, an inflation of roughly a thousandfold.
Editorial extensions
If this is right
- If the paradigm is correct, real-time interactive AI services will be built as two-model systems: on-device models absorb latency and personalization, while cloud models absorb complexity.
- The taxonomy predicts that a collaboration design is fully specified only when it says what is exchanged (data, features, or parameters), when the exchange happens, and whether interaction is single-device-to-cloud or multi-device-to-cloud.
- Systems must be co-designed with algorithms: data-based collaboration demands retrieval, indexing, and privacy tooling, while parameter-based collaboration demands on-device training support.
- Existing benchmarks with natural user-level partition, plus user-weighted metrics, become the default evaluation setup for hybrid learning, and new generative-task datasets with role-level partition extend them.
- Industrial deployments already exist in recommender systems, livestreaming understanding, and personal assistants, so the paradigm is not hypothetical.
Reading between the lines
- The taxonomy implies a testable design rule: the right exchange channel depends on the bottleneck being attacked -- data exchange targets personalization and privacy, feature exchange targets latency and bandwidth, and parameter exchange targets continual global improvement -- so future work could make this mapping explicit and quantitative.
- A natural extension the survey only gestures at is a three-layer cloud-edge-device architecture; if on-device models remain too weak and the cloud too far, medium-sized edge models would occupy the middle of this same taxonomy.
- Because the survey's quantitative motivation and some dataset statistics are not independently verified, a reader should treat the cost and scale numbers as directional until checked against primary sources; this is a verification task, not a reason to reject the taxonomy.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a survey of collaborative learning between on-device small models and cloud-based large models. It organizes the area into hardware, system/engine, algorithm, and application layers; proposes a data-based, feature-based, and parameter-based taxonomy of collaboration algorithms in Section 4.2.1; reviews representative academic and industrial advances; catalogs datasets and metrics; and closes with future directions. The central claim is that this paradigm addresses the latency, cost, personalization, and privacy bottlenecks of cloud-only large-model serving, and that the proposed taxonomy is a faithful organizing structure for the existing literature.
Significance. If corrected, the survey would be a genuinely useful reference: the layered framework is clear, the data/feature/parameter taxonomy is natural and well illustrated with concrete systems (e.g., EdgeRec/Taobao, Kuaishou, Apple Intelligence), and the breadth of coverage across more than 180 references, including system papers and industrial deployments, is a strength. Because the paper makes no technical derivations, its value rests on citation fidelity and internal consistency. The concrete errors identified below compromise that fidelity and must be fixed before the survey can be relied on as a reference work.
major comments (3)
- [Section 4.3, Table 2] The Reddit row in Table 2 lists 56,587,343 users for the LEAF Reddit dataset, but the cited LEAF benchmark [12] reports approximately 56,587 users. The sample count of 1,660,820 matches LEAF, so this appears to be a transcription error, but the published value is inflated by roughly 1000x and materially misrepresents a dataset that Table 2 uses as evidence of large-scale natural user-level partitions. This row must be corrected, and the remaining rows of Table 2 should be checked against their cited sources.
- [Section 1, cost estimate] The motivating cost estimate is internally inconsistent: 'approximately 30 million yuan per day' multiplied by 365 days is approximately 10.95 billion yuan, not 'over 9 billion yuan annually.' The sentence also attributes the estimate to Vivo while reference [75] is a Tencent Research Institute report, so the attribution is imprecise. Because this cost figure is the paper's primary quantitative motivation for the cost bottleneck, it should be corrected, verified against the cited source, and presented with a calculation the reader can reproduce.
- [Section 2.1, Table 1] Table 1 presents detailed hardware specifications for cloud servers and mobile devices without any source citations. Some entries are vendor-specific, generation-specific, and time-sensitive, and the table does not state whether numbers are peak or sustained values. For a survey whose framework includes a hardware layer, unsourced specification tables cannot be independently verified or updated. The authors should add explicit references or data-sheet links for each platform row and clarify the definition of each reported quantity.
minor comments (5)
- [Section 4.3, Table 2] The column layout of Table 2 is confusing: the 'Partition By' column contains user counts without a consistent unit label, and for the Reddit row the user count is placed where a partition descriptor is expected. Consider renaming the column and adding units consistently.
- [Figure 3] Q1.2 in Figure 3 reads 'university by masking hardware and software heterogeneity'; this should be 'universality.'
- [Figure 2] The memory/storage block in Figure 2 lists 'HHD'; this should be 'HDD.'
- [Section 4.2.2] The text contains a duplicated phrase, 'with with per-token confidence measure,' and Section 1 contains 'to to deliver efficient'; these should be cleaned up.
- [Sections 4.1-4.4] Several representative advances are drawn from the authors' own prior work (e.g., refs. [109, 123, 178, 161, 35, 55, 160]). This is understandable for a group central to the topic, but the authors should either broaden the representative examples or state their selection criteria so that the choice of examples is transparent.
Circularity Check
No circularity: survey makes no predictions or derivations; its paradigm claim and taxonomy rest on external literature rather than on the paper's own outputs.
full rationale
This is a survey paper, not a derivation, so the circularity patterns that apply to claimed predictions or first-principles derivations do not arise. The central claim that collaborative learning between on-device small models and cloud-based large models is a promising paradigm is supported by cited industrial reports and external systems (e.g., Qualcomm, Apple Intelligence, EdgeRec), not by the authors' own prior results. The data/feature/parameter taxonomy in Section 4.2.1 is a classification scheme imposed on existing work, not a quantity fitted from data and then re-predicted; categorizing methods by what is exchanged (raw samples, features, or parameters) is a descriptive organizational choice, not a circular derivation. The paper frequently cites the authors' own prior papers (e.g., refs. 33, 35, 55, 109, 123, 160, 161, 178) as representative advances, but these citations are used as examples of the surveyed literature, and the survey's framework does not reduce to the validity of any one of them. The manuscript does contain factual fidelity concerns, such as the inflated LEAF Reddit user count in Table 2 and the internally inconsistent cost estimate in Section 1, but those are correctness and citation-fidelity problems, not circular reasoning: the paper does not derive its conclusions from those numbers. No self-definitional reduction, no fitted-input-called-prediction step, and no load-bearing uniqueness argument imported from the authors' own prior work were found. Accordingly, the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Cited works and industry reports are summarized faithfully, with numbers transcribed correctly.
- domain assumption The four cloud bottlenecks (latency, cost, personalization, privacy) are the binding constraints, and the motivating cost estimate is accurate.
- domain assumption The selected systems, datasets, and deployments are representative of the field.
Cite this review
Pith. "Pith review of Collaborative Learning of On-Device Small Model and Cloud-Based Large Model: Advances and Future Directions." pith.science (2026). https://pith.science/paper/X54VX244
@misc{pith2026250415300,
author = {Pith},
title = {Pith review of: Collaborative Learning of On-Device Small Model and Cloud-Based Large Model: Advances and Future Directions},
year = {2026},
howpublished = {\url{https://pith.science/paper/X54VX244}},
note = {Machine review of arXiv:2504.15300}
}
read the original abstract
The conventional cloud-based large model learning framework is increasingly constrained by latency, cost, personalization, and privacy concerns. In this survey, we explore an emerging paradigm: collaborative learning between on-device small model and cloud-based large model, which promises low-latency, cost-efficient, and personalized intelligent services while preserving user privacy. We provide a comprehensive review across hardware, system, algorithm, and application layers. At each layer, we summarize key problems and recent advances from both academia and industry. In particular, we categorize collaboration algorithms into data-based, feature-based, and parameter-based frameworks. We also review publicly available datasets and evaluation metrics with user-level or device-level consideration tailored to collaborative learning settings. We further highlight real-world deployments, ranging from recommender systems and mobile livestreaming to personal intelligent assistants. We finally point out open research directions to guide future development in this rapidly evolving field.
Figures
Forward citations
Cited by 1 Pith paper
-
Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
A survey that builds a taxonomy of edge-cloud LLM-SLM collaboration for inference and training, claiming to be the first to unify both phases.
Reference graph
Works this paper leans on
-
[12]
Brendan McMahan, Virginia Smith, and Ameet Talwalkar
Sebastian Caldas, Peter Wu, Tian Li, Jakub Konečný, H. Brendan McMahan, Virginia Smith, and Ameet Talwalkar
-
[75]
Tencent Research Institute. 2024. The Surge of On-Device Large Models: Trends, Impacts, and Recommendations. https://www.tisi.org/30767/
2024
-
[1]
Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghe- mawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng
-
[2]
Alibaba Damo Academy. 2022. Top Ten Technology Trends of Damo Academy. https://damo.alibaba.com/events/ 32023091416946864330668487
2022
-
[3]
Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. 2024. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Vienna, Austria, 1–18. Collaborative Learning of On-D...
2024
-
[4]
Pranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju, Swaroop Mishra, Pei Zhou, Aditya Gupta, Dheeraj Rajagopal, Karthik Kappaganthu, Yiming Yang, Shyam Upadhyay, Manaal Faruqui, and Mausam
-
[5]
Samiul Alam, Luyang Liu, Ming Yan, and Mi Zhang. 2022. FedRolex: Model-Heterogeneous Federated Learning with Rolling Sub-Model Extraction. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., New Orleans, LA, USA, 14 pages
2022
-
[6]
Alimama. 2018. Ali_Display_Ad_Click. https://tianchi.aliyun.com/dataset/dataDetail?dataId=56
2018
Show all 187 references
-
[7]
Apple. 2024. Apple Intelligence. AI for the rest of us. https://www.apple.com/apple-intelligence/
2024
-
[8]
Apple. 2024. Apple Intelligence Foundation Language Models. CoRR abs/2407.21075 (2024), 47 pages
2024
-
[9]
Kallista Bonawitz, Hubert Eichner, Wolfgang Grieskamp, Dzmitry Huba, Alex Ingerman, Vladimir Ivanov, Chloé Kiddon, Jakub Konečný, Stefano Mazzocchi, Brendan McMahan, Timon Van Overveldt, David Petrou, Daniel Ramage, and Jason Roselander. 2019. Towards Federated Learning at Sca...
2019
-
[10]
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeffrey Wu. 2024. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. In...
2024
-
[11]
Han Cai, Chuang Gan, Ligeng Zhu, and Song Han. 2020. TinyTL: Reduce Memory, Not Parameters for Efficient On-Device Learning. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Virtual, 13 pages
2020
-
[13]
Andersen, Michael Kaminsky, and Subramanya Dulloor
Christopher Canel, Thomas Kim, Giulio Zhou, Conglong Li, Hyeontaek Lim, David G. Andersen, Michael Kaminsky, and Subramanya Dulloor. 2019. Scaling Video Analytics on Constrained Edge Nodes. In Proceedings of Machine Learning and Systems (MLSys) . mlsys.org, Stanford, Californi...
2019
-
[14]
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. 2023. Accelerating Large Language Model Decoding with Speculative Sampling. CoRR abs/2302.01318 (2023), 11 pages
2023 arXiv
-
[15]
Daoyuan Chen, Dawei Gao, Weirui Kuang, Yaliang Li, and Bolin Ding. 2022. pFL-Bench: A Comprehensive Benchmark for Personalized Federated Learning. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., New Orleans, LA, ...
2022
-
[16]
Fei Chen, Mi Luo, Zhenhua Dong, Zhenguo Li, and Xiuqiang He. 2019. Federated Meta-Learning with Fast Convergence and Efficient Communication. CoRR abs/1802.07876 (2019), 14 pages
2019 arXiv
-
[17]
Lingjiao Chen, Matei Zaharia, and James Zou. 2023. FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. CoRR abs/2305.05176 (2023), 13 pages
2023 arXiv
-
[18]
Kwok, and Yu Zhang
Shuhao Chen, Weisen Jiang, Baijiong Lin, James T. Kwok, and Yu Zhang. 2024. RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associate...
2024
-
[19]
Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Q. Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In Proceedings of USENIX S...
2018
-
[20]
Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Chen Qian, Yujia Qin, Xin Cong, Ruobing Xie, Zhiyuan Liu, Maosong Sun, and Jie Zhou. 2024. AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent B...
2024
-
[21]
Gonzalez, Ion Stoica, and Eric P
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. https://lmsys.org/blog/2023-...
2023
-
[22]
Yae Jee Cho, Luyang Liu, Zheng Xu, Aldi Fahrezi, and Gauri Joshi. 2024. Heterogeneous LoRA for Federated Fine- tuning of On-Device Foundation Models. In Proceedings of Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguis...
2024
-
[23]
Byung-Gon Chun, Sunghwan Ihm, Petros Maniatis, Mayur Naik, and Ashwin Patti. 2011. CloneCloud: elastic execution between mobile device and cloud. In Proceedings of European conference on Computer systems (EuroSys) . ACM, Salzburg, Austria, 301–314
2011
-
[24]
Kamil Ciosek. 2022. Imitation Learning by Reinforcement Learning. In Proceedings of International Conference on Learning Representations (ICLR). OpenReview.net, Virtual, 1–15. 24 Niu et al
2022
-
[25]
Liam Collins, Hamed Hassani, Aryan Mokhtari, and Sanjay Shakkottai. 2021. Exploiting Shared Representations for Personalized Federated Learning. In Proceedings of International Conference on Machine Learning (ICML) . PMLR, Virtual, 2089–2099
2021
-
[26]
Eduardo Cuervo, Aruna Balasubramanian, Dae-ki Cho, Alec Wolman, Stefan Saroiu, Ranveer Chandra, and Paramvir Bahl. 2010. MAUI: making smartphones last longer with code offload. In Proceedings of International Conference on Mobile Systems, Applications, and Services (MobiSys) ....
2010
-
[27]
Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAttention: Fast and Memory- Efficient Exact Attention with IO-Awareness. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., New Orlean...
2022
-
[28]
Robert David, Jared Duke, Advait Jain, Vijay Janapa Reddi, Nat Jeffries, Jian Li, Nick Kreeger, Ian Nappier, Meghna Natraj, Tiezhen Wang, Pete Warden, and Rocky Rhodes. 2021. TensorFlow Lite Micro: Embedded Machine Learning on TinyML Systems. InProceedings of Conference on Mac...
2021
-
[29]
DeepSeek-AI. 2025. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. CoRR abs/2501.12948 (2025), 22 pages
2025 arXiv
-
[30]
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., New Orleans, LA, USA, 15 pages
2022
-
[31]
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. QLoRA: Efficient Finetuning of Quantized LLMs. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., New Orleans, LA, USA, 28 pages
2023
-
[32]
Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Rühle, Laks V. S. Lakshmanan, and Ahmed Hassan Awadallah. 2024. Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing. In Proceedings of International Conference on Learning Representations (IC...
2024
-
[33]
Yucheng Ding, Chaoyue Niu, Fan Wu, Shaojie Tang, Chengfei Lyu, and Guihai Chen. 2022. On-Device Model Fine-Tuning with Label Correction in Recommender Systems. CoRR abs/2211.01163 (2022), 8 pages
2022 arXiv
-
[34]
Yucheng Ding, Chaoyue Niu, Fan Wu, Shaojie Tang, Chengfei Lyu, and Guihai Chen. 2023. DC-CCL: Device- Cloud Collaborative Controlled Learning for Large Vision Models. CoRR abs/2303.10361 (2023), 12 pages. https: //arxiv.org/abs/2303.10361
2023 arXiv
-
[35]
Yucheng Ding, Chaoyue Niu, Fan Wu, Shaojie Tang, Chengfei Lyu, and Guihai Chen. 2024. Enhancing On-Device LLM Inference with Historical Cloud-Based LLM Interactions. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) . ACM, Barcelona, Spain, 597–608
2024
-
[36]
Yucheng Ding, Chaoyue Niu, Fan Wu, Shaojie Tang, Chengfei Lyu, Yanghe Feng, and Guihai Chen. 2022. Federated Submodel Optimization for Hot and Cold Data Features. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., N...
2022
-
[37]
Yucheng Ding, Yangwenjian Tan, Xiangyu Liu, Chaoyue Niu, Fandong Meng, Jie Zhou, Ning Liu, Fan Wu, and Guihai Chen. 2025. Personalized Language Model Learning on Text Data Without User Identifiers. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and Data Mining ...
2025
-
[38]
Mahoney, and Kurt Keutzer
Zhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. 2019. HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-Precision. In Proceedings of IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, Seoul, Korea (South), 293–302
2019
-
[39]
Kuntai Du, Ahsan Pervaiz, Xin Yuan, Aakanksha Chowdhery, Qizheng Zhang, Henry Hoffmann, and Junchen Jiang
-
[40]
Ozdaglar
Alireza Fallah, Aryan Mokhtari, and Asuman E. Ozdaglar. 2020. Personalized Federated Learning with Theoretical Guarantees: A Model-Agnostic Meta-Learning Approach. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., ...
2020
-
[41]
Tiantian Feng, Digbalay Bose, Tuo Zhang, Rajat Hebbar, Anil Ramakrishna, Rahul Gupta, Mi Zhang, Salman Aves- timehr, and Shrikanth Narayanan. 2023. FedMultimodal: A Benchmark for Multimodal Federated Learning. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and ...
2023
-
[42]
Raphael Fischer and Amal Saadallah. 2024. AutoXPCR: Automated Multi-Objective Model Selection for Time Series Forecasting. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) . ACM, Barcelona, Spain, 806–815
2024
-
[43]
Jonathan Frankle and Michael Carbin. 2019. The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, New Orleans, LA, USA, 42 pages. Collaborative Learning of On-Dev...
2019
-
[44]
Elias Frantar and Dan Alistarh. 2023. SparseGPT: Massive Language Models Can be Accurately Pruned in One-Shot. In Proceedings of International Conference on Machine Learning (ICML) , Vol. 202. PMLR, Honolulu, Hawaii, USA, 10323–10337
2023
-
[45]
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. CoRR abs/2210.17323 (2022), 16 pages
2022 arXiv
-
[46]
Jingsheng Gao, Yixin Lian, Ziyi Zhou, Yuzhuo Fu, and Baoyuan Wang. 2023. LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming. InProceedings of Annual Meeting of the Association for Computational Linguistics (ACL). Association for...
2023
-
[47]
Georgi Gerganov. 2023. llama.cpp. https://github.com/ggerganov/llama.cpp
2023
-
[48]
Chen Gong, Zhenzhe Zheng, Fan Wu, Xiaofeng Jia, and Guihai Chen. 2024. Delta: A Cloud-assisted Data Enrichment Framework for On-Device Continual Learning. InProceedings of Annual International Conference on Mobile Computing and Networking (MobiCom). ACM, Washington D.C., DC, U...
2024
-
[49]
Xudong Gong, Qinlin Feng, Yuan Zhang, Jiangling Qin, Weijie Ding, Biao Li, Peng Jiang, and Kun Gai. 2022. Real-time Short Video Recommendation on Mobile Devices. In Proceedings of ACM International Conference on Information & Knowledge Management (CIKM). ACM, Atlanta, GA, USA,...
2022
-
[50]
Yu Gong, Ziwen Jiang, Yufei Feng, Binbin Hu, Kaiqi Zhao, Qingwen Liu, and Wenwu Ou. 2020. EdgeRec: Recommender System on Edge in Mobile Taobao. In Proceedings of ACM International Conference on Information and Knowledge Management (CIKM). ACM, Virtual, 2477–2484
2020
-
[51]
Google. 2017. TensorFlow Lite or LiteRT. https://ai.google.dev/edge/litert
2017
-
[52]
Google. 2024. MediaPipe LLM Inference API. https://ai.google.dev/edge/mediapipe/solutions/genai/llm_inference
2024
-
[53]
Gordon, David Ke Hong, Peter M
Mark S. Gordon, David Ke Hong, Peter M. Chen, Jason Flinn, Scott A. Mahlke, and Zhuoqing Morley Mao. 2015. Accelerating Mobile Applications through Flip-Flop Replication. In Proceedings of Annual International Conference on Mobile Systems, Applications, and Services (MobiSys) ...
2015
-
[54]
Gordon, Davoud Anoushe Jamshidi, Scott A
Mark S. Gordon, Davoud Anoushe Jamshidi, Scott A. Mahlke, Zhuoqing Morley Mao, and Xu Chen. 2012. COMET: Code Offload by Migrating Execution Transparently. In Proceedings of USENIX Symposium on Operating Systems Design and Implementation (OSDI) . USENIX Association, Hollywood,...
2012
-
[55]
Renjie Gu, Chaoyue Niu, Yikai Yan, Fan Wu, Shaojie Tang, Rongfeng Jia, Chengfei Lv, and Guihai Chen. 2022. On- Device Learning with Cloud-Coordinated Data Augmentation for Extreme Model Personalization in Recommender Systems. CoRR abs/2201.10382 (2022), 14 pages. https://arxiv...
2022 arXiv
-
[56]
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2024. MiniLLM: Knowledge Distillation of Large Language Models. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Vienna, Austria, 1–24
2024
-
[57]
Yuchao Gu, Xintao Wang, Jay Zhangjie Wu, Yujun Shi, Yunpeng Chen, Zihan Fan, Wuyou Xiao, Rui Zhao, Shuning Chang, Weijia Wu, Yixiao Ge, Ying Shan, and Mike Zheng Shou. 2023. Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion Models. In ...
2023
-
[58]
Otkrist Gupta and Ramesh Raskar. 2018. Distributed learning of deep neural network over multiple agents. Journal of Network and Computer Applications 116 (2018), 1–8
2018
-
[59]
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan. 2015. Deep Learning with Limited Numerical Precision. In Proceedings of the 32nd International Conference on Machine Learning (ICML) . JMLR.org, Lille, France, 1737–1746
2015
-
[60]
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. Retrieval Augmented Language Model Pre-Training. InProceedings of International Conference on Machine Learning (ICML). PMLR, Virtual, 3929–3938
2020
-
[61]
Song Han, Huizi Mao, and William J. Dally. 2016. Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding. In Proceedings of International Conference on Learning Representations, (ICLR). arXiv, San Juan, Puerto Rico, 14 pages
2016
-
[62]
Song Han, Jeff Pool, John Tran, and William J. Dally. 2015. Learning both Weights and Connections for Efficient Neural Networks. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Montreal, Quebec, Canada, 1135–1143
2015
-
[63]
Seungyeop Han, Haichen Shen, Matthai Philipose, Sharad Agarwal, Alec Wolman, and Arvind Krishnamurthy. 2016. MCDNN: An Approximation-Based Execution Framework for Deep Stream Processing Under Resource Constraints. In Proceedings of Annual International Conference on Mobile Sys...
2016
-
[64]
Hinton, Oriol Vinyals, and Jeffrey Dean
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015. Distilling the Knowledge in a Neural Network. CoRR abs/1503.02531 (2015), 9 pages. http://arxiv.org/abs/1503.02531
2015 arXiv
-
[65]
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-Efficient Transfer Learning for NLP. InProceedings of International Conference on Machine Learning (ICML) . PMLR, Lon...
2019
-
[66]
Le, Mark Sandler, Bo Chen, Weijun Wang, Liang-Chieh Chen, Mingxing Tan, Grace Chu, Vijay Vasudevan, and Yukun Zhu
Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le, Mark Sandler, Bo Chen, Weijun Wang, Liang-Chieh Chen, Mingxing Tan, Grace Chu, Vijay Vasudevan, and Yukun Zhu. 2019. Searching for MobileNetV3. In Proceedings of IEEE/CVF International Conference on Computer Vision (ICCV) ...
2019
-
[67]
Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. CoRR abs/1704.04861 (2017), 9 pages. https://arxiv.org/abs/1704.04861
2017 arXiv
-
[68]
Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. 2023. Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes. In Findings of the As...
2023
-
[69]
Tzu-Ming Harry Hsu, Hang Qi, and Matthew Brown. 2020. Federated Visual Classification with Real-World Data Distribution. In Proceedings of European Conference on Computer Vision (ECCV) . Springer, Glasgow, UK, 76–92
2020
-
[70]
Bo Hu and Wenjun Hu. 2019. LinkShare: device-centric control for concurrent and continuous mobile-cloud interactions. In Proceedings of ACM/IEEE Symposium on Edge Computing (SEC) . ACM, Arlington, Virginia, USA, 15–29
2019
-
[71]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
-
[72]
Kai Huang and Wei Gao. 2022. Real-time neural network inference on extremely weak devices: agile offloading with explainable AI. In Proceedings of Annual International Conference on Mobile Computing and Networking (MobiCom) . ACM, Sydney, NSW, Australia, 200–213
2022
-
[73]
Wenke Huang, Mang Ye, Zekun Shi, He Li, and Bo Du. 2023. Rethinking Federated Learning with Domain Shift: A Prototype View. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, Vancouver, BC, Canada, 16312–16322
2023
-
[74]
Hugging Face. 2018. Transformers. https://github.com/huggingface/transformers
2018
-
[76]
Howard, Hartwig Adam, and Dmitry Kalenichenko
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew G. Howard, Hartwig Adam, and Dmitry Kalenichenko. 2018. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. In Proceedings of IEEE Conference on Computer Visi...
2018
-
[77]
Chen Jia. 2024. Adversarial Moment-Matching Distillation of Large Language Models. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Vancouver, BC, Canada, 1–33
2024
-
[78]
Bowen Jin, Hansi Zeng, Zhenrui Yue, Dong Wang, Hamed Zamani, and Jiawei Han. 2025. Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning. CoRR abs/2503.09516 (2025), 16 pages
2025 arXiv
-
[79]
Whatmough, and Venkatesh Saligrama
Anil Kag, Igor Fedorov, Aditya Gangrade, Paul N. Whatmough, and Venkatesh Saligrama. 2023. Efficient Edge Inference by Selective Query. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Kigali, Rwanda, 25 pages
2023
-
[80]
Mudge, Jason Mars, and Lingjia Tang
Yiping Kang, Johann Hauswald, Cao Gao, Austin Rovinski, Trevor N. Mudge, Jason Mars, and Lingjia Tang. 2017. Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge. In Proceedings of International Conference on Architectural Support for Programming Language...
2017
-
[81]
Reddi, Sebastian U
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi, Sebastian U. Stich, and Ananda Theertha Suresh. 2020. SCAFFOLD: Stochastic Controlled Averaging for Federated Learning. In Proceedings of International Conference on Machine Learning (ICML) , Vol. 119. PML...
2020
-
[82]
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020. Generalization through Mem- orization: Nearest Neighbor Language Models. In Proceedings of International Conference on Learning Representations (ICLR). OpenReview.net, Addis Ababa, Ethiopia, 13 pages
2020
-
[83]
Yoon Kim and Alexander M. Rush. 2016. Sequence-Level Knowledge Distillation. In Proceedings of Conference on Empirical Methods in Natural Language Processing (EMNLP) . The Association for Computational Linguistics, Austin, Texas, USA, 1317–1327
2016
-
[84]
Jongwoo Ko, Sungnyun Kim, Tianyi Chen, and Se-Young Yun. 2024. DistiLLM: Towards Streamlined Distillation for Large Language Models. In Proceedings of International Conference on Machine Learning (ICML) . OpenReview.net, Vienna, Austria, 1–24
2024
-
[85]
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian U. Stich. 2020. A Unified Theory of Decentralized SGD with Changing Topology and Local Updates. In Proceedings of International Conference on Machine Learning (ICML). PMLR, Virtual, 5381–5393. Coll...
2020
-
[86]
Venieris, Mário Almeida, Ilias Leontiadis, and Nicholas D
Stefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis, and Nicholas D. Lane. 2020. SPINN: synergistic progressive inference of neural networks over device and cloud. In Proceedings of Annual International Conference on Mobile Computing and Networking (Mob...
2020
-
[87]
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The Power of Scale for Parameter-Efficient Prompt Tun- ing. In Proceedings of Conference on Empirical Methods in Natural Language Processing (EMNLP) . Association for Computational Linguistics, Punta Cana, Dominican Republic...
2021
-
[88]
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast Inference from Transformers via Speculative Decoding. In International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 202) . Proceedings of Machine Learning Research (PMLR), Hono...
2023
-
[89]
Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. InProceed...
2020
-
[90]
Ang Li, Jingwei Sun, Pengcheng Li, Yu Pu, Hai Li, and Yiran Chen. 2021. Hermes: an efficient federated learning framework for heterogeneous mobile clients. In Proceedings of Annual International Conference on Mobile Computing and Networking (MobiCom). ACM, New Orleans, Louisia...
2021
-
[91]
Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023. CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, I...
2023
-
[92]
Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren, Kai-Wei Chang, and Yejin Choi. 2023. Symbolic Chain-of- Thought Distillation: Small Models Can Also "Think" Step-by-Step. InProceedings of Annual Meeting of the Association for Computational Linguistics (ACL). Association...
2023
-
[93]
Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith. 2020. Federated Optimization in Heterogeneous Networks. In Proceedings of Conference on Machine Learning and Systems (MLSys) . mlsys.org, Austin, TX, USA, 22 pages
2020
-
[94]
Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Zhihua Zhang. 2020. On the Convergence of FedAvg on Non-IID Data. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Addis Ababa, Ethiopia, 26 pages
2020
-
[95]
Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2023. Contrastive Decoding: Open-ended Text Generation as Optimization. In Proceedings of Annual Meeting of the Association for Computational Linguistic...
2023
-
[96]
Xiang Lisa Li and Percy Liang. 2021. Prefix-Tuning: Optimizing Continuous Prompts for Generation. In Proceedings of Annual Meeting of the Association for Computational Linguistics (ACL) . ACL, Virtual, 4582–4597
2021
-
[97]
Yuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Yufei Wang, Guoqing Harry Xu, and Ravi Netravali. 2020. Reducto: On-camera Filtering for Resource-Efficient Real-Time Video Analytics. InProceedings of the ACM Special Interest Group on Data Communication on the applications, techno...
2020
-
[98]
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset. In Proceedings of International Joint Conference on Natural Language Processing (IJCNLP) . Asian Federation of Natural Language Processing...
2017
-
[99]
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2024. AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration. In Proceedings of the Seventh Annual Conference ...
2024
-
[100]
Stich, and Martin Jaggi
Tao Lin, Lingjing Kong, Sebastian U. Stich, and Martin Jaggi. 2020. Ensemble Distillation for Robust Model Fusion in Federated Learning. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., virtual, 13 pages
2020
-
[101]
Alisa Liu, Xiaochuang Han, Yizhong Wang, Yulia Tsvetkov, Yejin Choi, and Noah A. Smith. 2024. Tuning Lan- guage Models by Proxy. In Proceedings of Conference on Language Modeling (COLM) . OpenReview.net, Philadelphia, Pennsylvania, USA, 24 pages
2024
-
[102]
Jiaheng Liu, Chenchen Zhang, Jinyang Guo, Yuanxing Zhang, Haoran Que, Ken Deng, Zhiqi Bai, Jie Liu, Ge Zhang, Jiakai Wang, Yanan Wu, Congnan Liu, Jiamang Wang, Lin Qu, Wenbo Su, and Bo Zheng. 2024. DDK: Distilling Domain Knowledge for Efficient Large Language Models. In Procee...
2024
-
[103]
Yuejiang Liu and Alexandre Alahi. 2024. Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts. CoRR abs/2402.15505 (2024), 9 pages. 28 Niu et al
2024 arXiv
-
[104]
Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. 2017. Learning Efficient Convolutional Networks through Network Slimming. In Proceedings of IEEE International Conference on Computer Vision (ICCV). IEEE Computer Society, Venice, Italy, 2755–2763
2017
-
[105]
Zechun Liu, Barlas Oguz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, and Vikas Chandra. 2024. LLM-QAT: Data-Free Quantization Aware Training for Large Language Models. In Findings of the Association for Computational Ling...
2024
-
[106]
Llama Team, AI @ Meta. 2024. The Llama 3 Herd of Models. CoRR abs/2407.21783 (2024), 92 pages. https: //arxiv.org/abs/2407.21783
2024 arXiv
-
[107]
Keming Lu, Hongyi Yuan, Runji Lin, Junyang Lin, Zheng Yuan, Chang Zhou, and Jingren Zhou. 2024. Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models. In Proceedings of Conference of the North American Chapter of the Association for Computational Lin...
2024
-
[108]
Yan Lu, Yuanchao Shu, Xu Tan, Yunxin Liu, Mengyu Zhou, Qi Chen, and Dan Pei. 2019. Collaborative learning between cloud and end devices: an empirical study on location prediction. InProceedings of the ACM/IEEE Symposium on Edge Computing (SEC) . ACM/IEEE, Washington DC, USA, 139–151
2019
-
[109]
Chengfei Lv, Chaoyue Niu, Renjie Gu, Xiaotang Jiang, Zhaode Wang, Bin Liu, Ziqi Wu, Qiulin Yao, Congyu Huang, Panos Huang, Tao Huang, Hui Shu, Jinde Song, Bin Zou, Peng Lan, Guohuan Xu, Fei Wu, Shaojie Tang, Fan Wu, and Guihai Chen. 2022. Walle: An End-to-End, General-Purpose,...
2022
-
[110]
Zheqi Lv, Tianyu Zhan, Wenjie Wang, Xinyu Lin, Shengyu Zhang, Wenqiao Zhang, Jiwei Li, Kun Kuang, and Fei Wu. 2025. Collaboration of Large Language Models and Small Recommendation Models for Device-Cloud Recommendation. In Proceedings of ACM SIGKDD Conference on Knowledge Disc...
2025
-
[111]
Zheqi Lv, Wenqiao Zhang, Shengyu Zhang, Kun Kuang, Feng Wang, Yongwei Wang, Zhengyu Chen, Tao Shen, Hongxia Yang, Beng Chin Ooi, and Fei Wu. 2023. DUET: A Tuning-Free Device-Cloud Collaborative Parameters Generation Framework for Efficient Device Model Generalization. In Proce...
2023
-
[112]
Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2023. LLM-Pruner: On the Structural Pruning of Large Language Models. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., New Orleans, LA, USA, 19 pages
2023
-
[113]
Xiaosong Ma, Jie Zhang, Song Guo, and Wenchao Xu. 2022. Layer-wised Model Aggregation for Personalized Federated Learning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, New Orleans, LA, USA, 10082–10091
2022
-
[114]
Yishay Mansour, Mehryar Mohri, Jae Ro, and Ananda Theertha Suresh. 2020. Three Approaches for Personalization with Applications to Federated Learning. CoRR abs/2002.10619 (2020), 26 pages
2020 arXiv
-
[115]
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas. 2017. Communication- Efficient Learning of Deep Networks from Decentralized Data. In Proceedings of International Conference on Artificial Intelligence and Statistics (AISTATS). PMLR, Fort La...
2017
-
[116]
Meta. 2019. PyTorch Mobile. https://pytorch.org/mobile/home/
2019
-
[117]
Meta. 2023. ExecuTorch. https://github.com/pytorch/executorch
2023
-
[118]
Meta. 2024. torchchat. https://github.com/pytorch/torchchat
2024
-
[119]
Microsoft. 2020. DeepSpeed. https://github.com/microsoft/DeepSpeed
2020
-
[120]
Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius
Asit K. Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius. 2021. Accelerating Sparse Deep Neural Networks. CoRR abs/2104.08378 (2021), 18 pages
2021 arXiv
-
[121]
Eric Mitchell, Rafael Rafailov, Archit Sharma, Chelsea Finn, and Christopher D. Manning. 2024. An Emulator for Fine-tuning Large Language Models using Small Language Models. In Proceedings of International Conference on Learning Representations (ICLR). OpenReview.net, Vienna, ...
2024
-
[122]
Arindam Mitra, Hamed Khanpour, Corby Rosset, and Ahmed Awadallah. 2024. Orca-Math: Unlocking the potential of SLMs in Grade School Math. CoRR abs/2402.14830 (2024), 14 pages
2024 arXiv
-
[123]
Chaoyue Niu, Fan Wu, Shaojie Tang, Lifeng Hua, Rongfei Jia, Chengfei Lv, Zhihua Wu, and Guihai Chen. 2020. Billion-scale federated learning on mobile clients: a submodel design with tunable privacy. In Proceedings of Annual International Conference on Mobile Computing and Netw...
2020
-
[124]
NVIDIA. 2019. Megatron-LM. https://github.com/NVIDIA/Megatron-LM
2019
-
[125]
Gonzalez, M Waleed Kadous, and Ion Stoica
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, and Ion Stoica. 2025. RouteLLM: Learning to Route LLMs from Preference Data. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net,...
2025
-
[126]
Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, ...
2019
-
[127]
Ryan Po, Guandao Yang, Kfir Aberman, and Gordon Wetzstein. 2024. Orthogonal Adaptation for Modular Customiza- tion of Diffusion Models. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, Seattle, WA, USA, 7964–7973
2024
-
[128]
Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2024. ChatDev: Communicative Agents for Software Development. In Proceedings of Annual Meeting of the Associatio...
2024
-
[129]
Qualcomm. 2024. The future of AI is hybrid; Part I: Unlocking the generative AI future with on-device and hybrid AI. https://www.qualcomm.com/content/dam/qcomm-martech/dm-assets/documents/Whitepaper-The-future-of- AI-is-hybrid-Part-1-Unlocking-the-generative-AI-future-with-on-...
2024
-
[130]
Qualcomm. 2024. The future of AI is hybrid; Part II: Qualcomm is uniquely positioned to scale hybrid AI. https://www.qualcomm.com/content/dam/qcomm-martech/dm-assets/documents/Whitepaper-The-future-of- AI-is-hybrid-Part-2-Qualcomm-is-uniquely-positioned-to-scale-hybrid-AI.pdf
2024
-
[131]
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. 2016. XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks. In Proceedings of European Conference on Computer Vision (ECCV) . Springer, Amsterdam, The Netherlands, 525–542
2016
-
[132]
Reddi, Aditya Krishna Menon, Rohan Anil, and Sanjiv Kumar
Ankit Singh Rawat, Veeranjaneyulu Sadhanala, Afshin Rostamizadeh, Ayan Chakrabarti, Wittawat Jitkrittum, Vladimir Feinberg, Seungyeon Kim, Hrayr Harutyunyan, Nikunj Saunshi, Zachary Nado, Rakesh Shivanna, Sashank J. Reddi, Aditya Krishna Menon, Rohan Anil, and Sanjiv Kumar. 20...
2024 arXiv
-
[133]
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2015. FitNets: Hints for Thin Deep Nets. CoRR abs/1412.6550 (2015), 13 pages
2015 arXiv
-
[134]
Yichen Ruan, Xiaoxi Zhang, Shu-Che Liang, and Carlee Joe-Wong. 2021. Towards Flexible Device Participation in Federated Learning. In Proceedings of International Conference on Artificial Intelligence and Statistics (AISTATS) . PMLR, Virtual, 3403–3411
2021
-
[135]
Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen
Mark Sandler, Andrew G. Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. 2018. MobileNetV2: Inverted Residuals and Linear Bottlenecks. InProceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Computer Vision Foundation / IEEE Computer So...
2018
-
[136]
Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Tran, Yi Tay, and Donald Metzler. 2022. Confident Adaptive Language Modeling. InProceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., New Orleans, LA...
2022
-
[137]
Virginia Smith, Chao-Kai Chiang, Maziar Sanjabi, and Ameet Talwalkar. 2017. Federated Multi-Task Learning. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Long Beach, CA, USA, 4424–4434
2017
-
[138]
Nimit Sharad Sohoni, Christopher Richard Aberger, Megan Leszczynski, Jian Zhang, and Christopher Ré. 2019. Low-Memory Neural Network Training: A Technical Report. CoRR abs/1904.10631 (2019), 38 pages
2019 arXiv
-
[139]
Nikita Starodubcev, Dmitry Baranchuk, Artem Fedorov, and Artem Babenko. 2024. Your Student is Better than Expected: Adaptive Teacher-Student Collaboration for Text-Conditional Diffusion Models. InProceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CV...
2024
-
[140]
Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V. Le. 2019. MnasNet: Platform-Aware Neural Architecture Search for Mobile. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Computer Vision Foundat...
2019
-
[141]
Mingxing Tan and Quoc V. Le. 2019. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of International Conference on Machine Learning (ICML) . PMLR, Long Beach, CA, USA, 6105–6114
2019
-
[142]
Yue Tan, Guodong Long, Lu Liu, Tianyi Zhou, Qinghua Lu, Jing Jiang, and Chengqi Zhang. 2022. FedProto: Federated Prototype Learning across Heterogeneous Clients. In Proceedings of AAAI Conference on Artificial Intelligence (AAAI) . AAAI Press, Virtual, 8432–8440
2022
-
[143]
Yue Tan, Guodong Long, Jie Ma, Lu Liu, Tianyi Zhou, and Jing Jiang. 2022. Federated Learning from Pre-Trained Models: A Contrastive Learning Approach. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., New Orleans, ...
2022
-
[144]
Hashimoto
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford Alpaca: An Instruction-following LLaMA model. https://github.com/tatsu-lab/stanford_ 30 Niu et al. alpaca
2023
-
[145]
MLC team. 2023. MLC-LLM. https://github.com/mlc-ai/mlc-llm
2023
-
[146]
Surat Teerapittayanon, Bradley McDanel, and H. T. Kung. 2016. BranchyNet: Fast inference via early exiting from deep neural networks. In Proceedings of International Conference on Pattern Recognition (ICPR) . IEEE, Cancún, Mexico, 2464–2469
2016
-
[147]
Surat Teerapittayanon, Bradley McDanel, and H. T. Kung. 2017. Distributed Deep Neural Networks Over the Cloud, the Edge and End Devices. In Proceedings of IEEE International Conference on Distributed Computing Systems (ICDCS) . IEEE, Atlanta, GA, USA, 328–339
2017
-
[148]
Vincent Poor
Jianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi, and H. Vincent Poor. 2020. Tackling the Objective Inconsistency Problem in Heterogeneous Federated Optimization. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., V...
2020
-
[149]
Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han. 2019. HAQ: Hardware-Aware Automated Quantization With Mixed Precision. InProceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Computer Vision Foundation / IEEE, Long Beach, CA, USA, 8612–8620
2019
-
[150]
Qipeng Wang, Mengwei Xu, Chao Jin, Xinran Dong, Jinliang Yuan, Xin Jin, Gang Huang, Yunxin Liu, and Xuanzhe Liu. 2022. Melon: breaking the memory wall for resource-efficient on-device machine learning. In Proceedings of Annual International Conference on Mobile Systems, Applic...
2022
-
[151]
Yiding Wang, Kai Chen, Haisheng Tan, and Kun Guo. 2023. Tabi: An Efficient Multi-Level Inference System for Large Language Models. In Proceedings of European Conference on Computer Systems (EuroSys) . ACM, Rome, Italy, 233–248
2023
-
[152]
Heming Xia, Tao Ge, Peiyi Wang, Si-Qing Chen, Furu Wei, and Zhifang Sui. 2023. Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation. InFindings of the Association for Computational Linguistics (EMNLP). Association for Computational Linguis...
2023
-
[153]
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen. 2024. Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning. InProceedings of The Twelfth International Conference on Learning Representations (ICLR). OpenReview.net, Vienna, Austria, 25 pages
2024
-
[154]
Guangxuan Xiao, Ji Lin, and Song Han. 2023. Offsite-Tuning: Transfer Learning without Full Model. CoRR abs/2302.04870 (2023), 12 pages. https://arxiv.org/abs/2302.04870
2023 arXiv
-
[155]
Guangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu, Julien Demouth, and Song Han. 2023. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models. In Proceedings of International Conference on Machine Learning (ICML), Vol. 202. PMLR, Honolulu, Hawaii...
2023
-
[156]
Hovy, and Quoc V
Qizhe Xie, Minh-Thang Luong, Eduard H. Hovy, and Quoc V. Le. 2020. Self-Training With Noisy Student Improves ImageNet Classification. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Computer Vision Foundation / IEEE, Seattle, WA, USA, ...
2020
- [157]
-
[158]
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang
-
[159]
Xinchao Xu, Zhibin Gou, Wenquan Wu, Zheng-Yu Niu, Hua Wu, Haifeng Wang, and Shihang Wang. 2022. Long Time No See! Open-Domain Conversation with Long-Term Persona Memory. In Findings of the Association for Computational Linguistics (ACL). Association for Computational Linguisti...
2022
-
[160]
Yikai Yan, Chaoyue Niu, Yucheng Ding, Zhenzhe Zheng, Shaojie Tang, Qinya Li, Fan Wu, Chengfei Lyu, Yanghe Feng, and Guihai Chen. 2024. Federated Optimization Under Intermittent Client Availability. INFORMS Journal on Computing 36, 1 (2024), 185–202
2024
-
[161]
Yikai Yan, Chaoyue Niu, Renjie Gu, Fan Wu, Shaojie Tang, Lifeng Hua, Chengfei Lyu, and Guihai Chen. 2022. On- Device Learning for Model Personalization with Large-Scale Cloud-Coordinated Domain Adaption. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and Data M...
2022
-
[162]
Jiangchao Yao, Feng Wang, Kunyang Jia, Bo Han, Jingren Zhou, and Hongxia Yang. 2021. Device-Cloud Collaborative Learning for Recommendation. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD). ACM, Virtual, 3865–3874
2021
-
[163]
Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. 2022. ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers. InProceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) ....
2022
-
[164]
In Proceedings of International Conference on Learning Representations (ICLR)
WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex Instructions. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Vienna, Austria, 1–22
-
[165]
Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2024. MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models. In Proceedings of International Conference on Learning Repre...
2024
-
[166]
Murong Yue, Jie Zhao, Min Zhang, Liang Du, and Ziyu Yao. 2024. Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient Reasoning. In Proceedings of International Conference on Learning Representations (ICLR). OpenReview.net, Vienna, Austria, 38 pages
2024
-
[167]
Kaiyan Zhang, Jianyu Wang, Ning Ding, Biqing Qi, Ermo Hua, Xingtai Lv, and Bowen Zhou. 2024. Fast and Slow Generating: An Empirical Study on Large and Small Language Models Collaborative Decoding. CoRR abs/2406.12295 (2024), 17 pages. https://arxiv.org/abs/2406.12295
2024 arXiv
-
[168]
Kaiyan Zhang, Jianyu Wang, Ermo Hua, Biqing Qi, Ning Ding, and Bowen Zhou. 2024. CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction Following. In Proceedings of Annual Meeting of the Association for Computational Linguisti...
2024
-
[169]
Michael Zhang, Karan Sapra, Sanja Fidler, Serena Yeung, and José M. Álvarez. 2021. Personalized Federated Learning with First Order Model Optimization. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Virtual, 17 pages
2021
-
[170]
Junho Yim, Donggyu Joo, Ji-Hoon Bae, and Junmo Kim. 2017. A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer Learning. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE Computer Society, Honolulu, HI...
2017
-
[171]
Hospedales, and Huchuan Lu
Ying Zhang, Tao Xiang, Timothy M. Hospedales, and Huchuan Lu. 2018. Deep Mutual Learning. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Computer Vision Foundation / IEEE Computer Society, Salt Lake City, UT, USA, 4320–4328
2018
-
[172]
Yinhe Zheng, Guanyi Chen, Minlie Huang, Song Liu, and Xuan Zhu. 2019. Personalized Dialogue Generation with Diversified Traits. CoRR abs/1901.09672 (2019), 12 pages
2019 arXiv
-
[173]
Qihuang Zhong, Liang Ding, Li Shen, Juhua Liu, Bo Du, and Dacheng Tao. 2024. Revisiting Knowledge Distillation for Autoregressive Language Models. In Proceedings of Annual Meeting of the Association for Computational Linguistics (ACL). Association for Computational Linguistics...
2024
-
[174]
Shuchang Zhou, Zekun Ni, Xinyu Zhou, He Wen, Yuxin Wu, and Yuheng Zou. 2016. DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients. CoRR abs/1606.06160 (2016), 13 pages
2016 arXiv
-
[175]
Zhanhui Zhou, Zhixuan Liu, Jie Liu, Zhichen Dong, Chao Yang, and Yu Qiao. 2024. Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associat...
2024
-
[176]
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Personalizing Dialogue Agents: I have a dog, do you have pets too?. In Proceedings of Annual Meeting of the Association for Computational Linguistics (ACL). Association for Computatio...
2018
-
[177]
Han Zhu, Junqi Jin, Chang Tan, Fei Pan, Yifan Zeng, Han Li, and Kun Gai. 2017. Optimized Cost per Click in Taobao Display Advertising. In Proceedings of ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD). ACM, Halifax, NS, Canada, 2191–2200
2017
-
[178]
Yufei Zhu, Chaoyue Niu, Yikai Yan, Zhijie Cao, Hao Jiang, Chengfei Lyu, Shaojie Tang, and Fan Wu. 2023. Device- Unimodal Cloud-Multimodal Collaboration for Livestreaming Content Understanding. In Proceedings of IEEE Interna- tional Conference on Data Mining (ICDM) . IEEE, Shan...
2023
-
[179]
Zhuangdi Zhu, Junyuan Hong, and Jiayu Zhou. 2021. Data-Free Knowledge Distillation for Heterogeneous Federated Learning. In Proceedings of International Conference on Machine Learning (ICML) . PMLR, Virtual, 12878–12889
2021
-
[180]
Richard Zhuang, Tianhao Wu, Zhaojin Wen, Andrew Li, Jiantao Jiao, and Kannan Ramchandran. 2025. EmbedLLM: Learning Compact Representations of Large Language Models. In Proceedings of International Conference on Learning Representations (ICLR). OpenReview.net, Singapore, 14 pages
2025
-
[181]
Barret Zoph and Quoc V. Le. 2017. Neural Architecture Search with Reinforcement Learning. In Proceedings of International Conference on Learning Representations, (ICLR) . OpenReview.net, Toulon, France, 16 pages
2017
-
[182]
Barrett, Michael I
Banghua Zhu, Ying Sheng, Lianmin Zheng, Clark W. Barrett, Michael I. Jordan, and Jiantao Jiao. 2023. On Optimal Caching and Model Selection for Large Model Inference. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc...
2023
-
[2016]
In Proceedings of USENIX Symposium on Operating Systems Design and Implementation (OSDI)
TensorFlow: A System for Large-Scale Machine Learning. In Proceedings of USENIX Symposium on Operating Systems Design and Implementation (OSDI) . USENIX, Savannah, GA, USA, 265–283
-
[2018]
CoRR abs/1812.01097 (2018), 9 pages
LEAF: A Benchmark for Federated Settings. CoRR abs/1812.01097 (2018), 9 pages
2018 arXiv
-
[2020]
In Proceedings of Annual Conference of ACM’s Special Interest Group on Data Communication (SIGCOMM)
Server-driven video streaming for deep learning inference. In Proceedings of Annual Conference of ACM’s Special Interest Group on Data Communication (SIGCOMM) . ACM, Virtual, 557–570
-
[2022]
In Proceedings of International Conference on Learning Representations (ICLR)
LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of International Conference on Learning Representations (ICLR). OpenReview.net, Virtual, 13 pages
-
[2024]
In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS)
AutoMix: Automatically Mixing Language Models. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., Vancouver, BC, Canada, 35 pages
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.