Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Collaborative Learning of On-Device Small Model and Cloud-Based Large Model: Advances and Future Directions

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This survey argues that cloud-only large-model learning has hit latency, cost, personalization, and privacy limits, and that on-device small models collaborating with cloud large models is the emerging fix.

desk verdict A capable survey with a genuinely useful taxonomy, marred by a 1000x dataset error and sloppy cost arithmetic that are fixable in review. read the letter →

arxiv 2504.15300 v1 pith:X54VX244 submitted 2025-04-17 cs.LG cs.DCcs.MA

classification cs.LGcs.DCcs.MA
keywords collaborativelearningon-devicesmallmodelcloud-basedlargehybridAIfederatedknowledgedistillationcompressionedgeintelligence
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that the standard cloud-only framework for large-model learning is hitting four binding constraints -- high response latency, high cost and cloud load, weak personalization, and privacy risk -- and that the emerging answer is collaborative learning between a small model on each device and a large model in the cloud. A sympathetic reader is being asked to accept that this hybrid paradigm is real, that it already has production deployments, and that the field's work can be usefully organized by a layered framework spanning hardware, system, algorithm, and application. The survey's central organizing device is a taxonomy of collaboration algorithms by what the two sides exchange: data, features, or parameters. If the survey is right, future AI services will increasingly be distributed systems that keep private data on-device, adapt to individual users locally, and call the cloud only for what the small model cannot do.

What carries the argument

The load-bearing organizing device is a layered framework -- hardware, system and engine, model and algorithm, application -- together with a taxonomy of collaboration algorithms based on what is exchanged between cloud and device: data-based collaboration (raw samples or queries, including data filtering and query routing), feature-based collaboration (intermediate or final model outputs, including model splitting, early exiting, distillation, and parallel decoding), and parameter-based collaboration (models or model updates, including federated learning, model ensemble, and offsite or proxy tuning). This taxonomy is what lets the survey map dozens of works onto a small number of design choices: where subtasks are allocated, when exchange happens, and what is transmitted. The framework's second function is to show that the four identified bottlenecks each correspond to a layer where a concrete problem and a line of advances already exist.

What would settle it

Check the primary source [12] for the Reddit dataset's user count; if it is about 56,587 rather than 56,587,343, the survey's Table 2 contains a thousandfold error, and the same check should be run on the other dataset rows and on the industry cost estimate cited as [75] (about 30 million yuan per day for 300 million users).

Watch

Extended reading notes

Core claim

The paper's central claim is that collaborative learning between an on-device small model and a cloud-based large model can escape four bottlenecks of cloud-centric learning: latency too high for real-time interaction, cost and load too high when millions of devices upload raw data, a single global model that cannot personalize to individual users, and privacy risk from centralizing sensitive data. In the proposed paradigm, the small model handles local real-time inference, adapts to the user's data, and sends only non-sensitive samples, features, or parameter updates upward, while the large model transfers knowledge down via distillation or compression and handles complex global reasoning. The survey claims this creates a virtuous co-evolution cycle and supports it with a four-layer framework review and with industrial deployments in recommender systems, livestreaming content understanding, and personal intelligent assistants. It also argues that existing benchmarks fall short because they lack user-level or device-level partitioning and metrics.

Load-bearing premise

The survey's whole edifice rests on faithfully summarizing more than 180 cited works and industry reports; one concrete failure is visible already, since Table 2 lists the Reddit dataset as partitioned across 56,587,343 users while the cited benchmark [12] reports about 56,587 users, an inflation of roughly a thousandfold.

Editorial extensions

If this is right

  • If the paradigm is correct, real-time interactive AI services will be built as two-model systems: on-device models absorb latency and personalization, while cloud models absorb complexity.
  • The taxonomy predicts that a collaboration design is fully specified only when it says what is exchanged (data, features, or parameters), when the exchange happens, and whether interaction is single-device-to-cloud or multi-device-to-cloud.
  • Systems must be co-designed with algorithms: data-based collaboration demands retrieval, indexing, and privacy tooling, while parameter-based collaboration demands on-device training support.
  • Existing benchmarks with natural user-level partition, plus user-weighted metrics, become the default evaluation setup for hybrid learning, and new generative-task datasets with role-level partition extend them.
  • Industrial deployments already exist in recommender systems, livestreaming understanding, and personal assistants, so the paradigm is not hypothetical.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy implies a testable design rule: the right exchange channel depends on the bottleneck being attacked -- data exchange targets personalization and privacy, feature exchange targets latency and bandwidth, and parameter exchange targets continual global improvement -- so future work could make this mapping explicit and quantitative.
  • A natural extension the survey only gestures at is a three-layer cloud-edge-device architecture; if on-device models remain too weak and the cloud too far, medium-sized edge models would occupy the middle of this same taxonomy.
  • Because the survey's quantitative motivation and some dataset statistics are not independently verified, a reader should treat the cost and scale numbers as directional until checked against primary sources; this is a verification task, not a reason to reject the taxonomy.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper is a survey of collaborative learning between on-device small models and cloud-based large models. It organizes the area into hardware, system/engine, algorithm, and application layers; proposes a data-based, feature-based, and parameter-based taxonomy of collaboration algorithms in Section 4.2.1; reviews representative academic and industrial advances; catalogs datasets and metrics; and closes with future directions. The central claim is that this paradigm addresses the latency, cost, personalization, and privacy bottlenecks of cloud-only large-model serving, and that the proposed taxonomy is a faithful organizing structure for the existing literature.

Significance. If corrected, the survey would be a genuinely useful reference: the layered framework is clear, the data/feature/parameter taxonomy is natural and well illustrated with concrete systems (e.g., EdgeRec/Taobao, Kuaishou, Apple Intelligence), and the breadth of coverage across more than 180 references, including system papers and industrial deployments, is a strength. Because the paper makes no technical derivations, its value rests on citation fidelity and internal consistency. The concrete errors identified below compromise that fidelity and must be fixed before the survey can be relied on as a reference work.

major comments (3)
  1. [Section 4.3, Table 2] The Reddit row in Table 2 lists 56,587,343 users for the LEAF Reddit dataset, but the cited LEAF benchmark [12] reports approximately 56,587 users. The sample count of 1,660,820 matches LEAF, so this appears to be a transcription error, but the published value is inflated by roughly 1000x and materially misrepresents a dataset that Table 2 uses as evidence of large-scale natural user-level partitions. This row must be corrected, and the remaining rows of Table 2 should be checked against their cited sources.
  2. [Section 1, cost estimate] The motivating cost estimate is internally inconsistent: 'approximately 30 million yuan per day' multiplied by 365 days is approximately 10.95 billion yuan, not 'over 9 billion yuan annually.' The sentence also attributes the estimate to Vivo while reference [75] is a Tencent Research Institute report, so the attribution is imprecise. Because this cost figure is the paper's primary quantitative motivation for the cost bottleneck, it should be corrected, verified against the cited source, and presented with a calculation the reader can reproduce.
  3. [Section 2.1, Table 1] Table 1 presents detailed hardware specifications for cloud servers and mobile devices without any source citations. Some entries are vendor-specific, generation-specific, and time-sensitive, and the table does not state whether numbers are peak or sustained values. For a survey whose framework includes a hardware layer, unsourced specification tables cannot be independently verified or updated. The authors should add explicit references or data-sheet links for each platform row and clarify the definition of each reported quantity.
minor comments (5)
  1. [Section 4.3, Table 2] The column layout of Table 2 is confusing: the 'Partition By' column contains user counts without a consistent unit label, and for the Reddit row the user count is placed where a partition descriptor is expected. Consider renaming the column and adding units consistently.
  2. [Figure 3] Q1.2 in Figure 3 reads 'university by masking hardware and software heterogeneity'; this should be 'universality.'
  3. [Figure 2] The memory/storage block in Figure 2 lists 'HHD'; this should be 'HDD.'
  4. [Section 4.2.2] The text contains a duplicated phrase, 'with with per-token confidence measure,' and Section 1 contains 'to to deliver efficient'; these should be cleaned up.
  5. [Sections 4.1-4.4] Several representative advances are drawn from the authors' own prior work (e.g., refs. [109, 123, 178, 161, 35, 55, 160]). This is understandable for a group central to the topic, but the authors should either broaden the representative examples or state their selection criteria so that the choice of examples is transparent.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: survey makes no predictions or derivations; its paradigm claim and taxonomy rest on external literature rather than on the paper's own outputs.

full rationale

This is a survey paper, not a derivation, so the circularity patterns that apply to claimed predictions or first-principles derivations do not arise. The central claim that collaborative learning between on-device small models and cloud-based large models is a promising paradigm is supported by cited industrial reports and external systems (e.g., Qualcomm, Apple Intelligence, EdgeRec), not by the authors' own prior results. The data/feature/parameter taxonomy in Section 4.2.1 is a classification scheme imposed on existing work, not a quantity fitted from data and then re-predicted; categorizing methods by what is exchanged (raw samples, features, or parameters) is a descriptive organizational choice, not a circular derivation. The paper frequently cites the authors' own prior papers (e.g., refs. 33, 35, 55, 109, 123, 160, 161, 178) as representative advances, but these citations are used as examples of the surveyed literature, and the survey's framework does not reduce to the validity of any one of them. The manuscript does contain factual fidelity concerns, such as the inflated LEAF Reddit user count in Table 2 and the internally inconsistent cost estimate in Section 1, but those are correctness and citation-fidelity problems, not circular reasoning: the paper does not derive its conclusions from those numbers. No self-definitional reduction, no fitted-input-called-prediction step, and no load-bearing uniqueness argument imported from the authors' own prior work were found. Accordingly, the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey introduces no free parameters and no new entities. Its central content is a recitation of the cited literature, so the load-bearing items are the fidelity of the summarization (violated at least once in Table 2) and the representativeness of the selection. These are audit-able in principle but not fully verified here.

assumptions (3)
  • domain assumption Cited works and industry reports are summarized faithfully, with numbers transcribed correctly.
    All survey claims about the state of the art depend on this. It is violated at least once: Table 2 lists the LEAF Reddit dataset with 56,587,343 users, roughly 1000x the value reported in the cited LEAF paper.
  • domain assumption The four cloud bottlenecks (latency, cost, personalization, privacy) are the binding constraints, and the motivating cost estimate is accurate.
    Section 1 motivates the entire paradigm from these bottlenecks and cites a single Tencent Research Institute report for the Vivo cost figure of about 30 million yuan per day; the survey does not verify the figure.
  • domain assumption The selected systems, datasets, and deployments are representative of the field.
    Sections 4.2 to 4.4 present representative advances and apps without an explicit selection criterion, and the authors' own publications occupy several representative slots (refs. 33, 35, 123, 178).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Collaborative Learning of On-Device Small Model and Cloud-Based Large Model: Advances and Future Directions." pith.science (2026). https://pith.science/paper/X54VX244

@misc{pith2026250415300,
  author       = {Pith},
  title        = {Pith review of: Collaborative Learning of On-Device Small Model and Cloud-Based Large Model: Advances and Future Directions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/X54VX244}},
  note         = {Machine review of arXiv:2504.15300}
}
read the original abstract

The conventional cloud-based large model learning framework is increasingly constrained by latency, cost, personalization, and privacy concerns. In this survey, we explore an emerging paradigm: collaborative learning between on-device small model and cloud-based large model, which promises low-latency, cost-efficient, and personalized intelligent services while preserving user privacy. We provide a comprehensive review across hardware, system, algorithm, and application layers. At each layer, we summarize key problems and recent advances from both academia and industry. In particular, we categorize collaboration algorithms into data-based, feature-based, and parameter-based frameworks. We also review publicly available datasets and evaluation metrics with user-level or device-level consideration tailored to collaborative learning settings. We further highlight real-world deployments, ranging from recommender systems and mobile livestreaming to personal intelligent assistants. We finally point out open research directions to guide future development in this rapidly evolving field.

Figures

Figures reproduced from arXiv: 2504.15300 by the authors.

Figure 1
Figure 1. Learning paradigms comparison. concurrent user requests, response latency increases significantly. In scenarios with poor mobile network conditions, cloud interface access limits, or cloud service failures, the cloud-based large model service becomes unavailable. The second bottleneck is high cost and heavy load. On the mobile device side, uploading raw data can result in significant cellular data usage, particularl… view at source ↗
Figure 2
Figure 2. Overall framework. varies from a few W to several tens of W. In contrast, microcontrollers have a power consumption ranging from hundreds of mW to a few W; and (4) from networking, smartphones and XR devices connect to WiFi and 3G/4G/5G networks, with download speeds typically surpassing upload speeds. For example, the Snapdragon X75 5G modem offers a downlink of up to 10 Gbps and an uplink of up to 3.5 Gbps. Embedd… view at source ↗
Figure 3
Figure 3. Key problems at different layers. assistant (e.g., Apple’s Siri, Amazon’s Alexa, Huawei’s Xiaoyi, Xiaomi’s Xiaoai, OPPO’s Xiaobu), and chatbot (e.g., OpenAI’s ChatGPT, Alibaba’s Qwen, Anthropic’s Claude, Google’s Gemini); (3) for user behavior data analysis scenarios, common applications include recommender systems for e-commerce platforms (e.g., Amazon, Taobao, Jingdong, Meituan), social media networks (e.g., Faceb… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges

    cs.DC 2025-07 conditional novelty 4.0 of 10

    A survey that builds a taxonomy of edge-cloud LLM-SLM collaboration for inference and training, claiming to be the first to unify both phases.

Reference graph

Works this paper leans on

187 extracted references · 53 canonical work pages · cited by 1 Pith paper

  1. [12]

    Brendan McMahan, Virginia Smith, and Ameet Talwalkar

    Sebastian Caldas, Peter Wu, Tian Li, Jakub Konečný, H. Brendan McMahan, Virginia Smith, and Ameet Talwalkar

  2. [75]

    Tencent Research Institute. 2024. The Surge of On-Device Large Models: Trends, Impacts, and Recommendations. https://www.tisi.org/30767/

  3. [1]

    Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng

    Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghe- mawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng

  4. [2]

    Alibaba Damo Academy. 2022. Top Ten Technology Trends of Damo Academy. https://damo.alibaba.com/events/ 32023091416946864330668487

  5. [3]

    Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. 2024. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Vienna, Austria, 1–18. Collaborative Learning of On-D...

  6. [4]

    Pranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju, Swaroop Mishra, Pei Zhou, Aditya Gupta, Dheeraj Rajagopal, Karthik Kappaganthu, Yiming Yang, Shyam Upadhyay, Manaal Faruqui, and Mausam

  7. [5]

    Samiul Alam, Luyang Liu, Ming Yan, and Mi Zhang. 2022. FedRolex: Model-Heterogeneous Federated Learning with Rolling Sub-Model Extraction. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., New Orleans, LA, USA, 14 pages

  8. [6]

    Alimama. 2018. Ali_Display_Ad_Click. https://tianchi.aliyun.com/dataset/dataDetail?dataId=56

Show all 187 references
  1. [7]

    Apple. 2024. Apple Intelligence. AI for the rest of us. https://www.apple.com/apple-intelligence/

  2. [8]

    Apple. 2024. Apple Intelligence Foundation Language Models. CoRR abs/2407.21075 (2024), 47 pages

  3. [9]

    Kallista Bonawitz, Hubert Eichner, Wolfgang Grieskamp, Dzmitry Huba, Alex Ingerman, Vladimir Ivanov, Chloé Kiddon, Jakub Konečný, Stefano Mazzocchi, Brendan McMahan, Timon Van Overveldt, David Petrou, Daniel Ramage, and Jason Roselander. 2019. Towards Federated Learning at Sca...

  4. [10]

    Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeffrey Wu. 2024. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. In...

  5. [11]

    Han Cai, Chuang Gan, Ligeng Zhu, and Song Han. 2020. TinyTL: Reduce Memory, Not Parameters for Efficient On-Device Learning. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Virtual, 13 pages

  6. [13]

    Andersen, Michael Kaminsky, and Subramanya Dulloor

    Christopher Canel, Thomas Kim, Giulio Zhou, Conglong Li, Hyeontaek Lim, David G. Andersen, Michael Kaminsky, and Subramanya Dulloor. 2019. Scaling Video Analytics on Constrained Edge Nodes. In Proceedings of Machine Learning and Systems (MLSys) . mlsys.org, Stanford, Californi...

  7. [14]

    Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. 2023. Accelerating Large Language Model Decoding with Speculative Sampling. CoRR abs/2302.01318 (2023), 11 pages

  8. [15]

    Daoyuan Chen, Dawei Gao, Weirui Kuang, Yaliang Li, and Bolin Ding. 2022. pFL-Bench: A Comprehensive Benchmark for Personalized Federated Learning. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., New Orleans, LA, ...

  9. [16]

    Fei Chen, Mi Luo, Zhenhua Dong, Zhenguo Li, and Xiuqiang He. 2019. Federated Meta-Learning with Fast Convergence and Efficient Communication. CoRR abs/1802.07876 (2019), 14 pages

  10. [17]

    Lingjiao Chen, Matei Zaharia, and James Zou. 2023. FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. CoRR abs/2305.05176 (2023), 13 pages

  11. [18]

    Kwok, and Yu Zhang

    Shuhao Chen, Weisen Jiang, Baijiong Lin, James T. Kwok, and Yu Zhang. 2024. RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associate...

  12. [19]

    Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy

    Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Q. Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In Proceedings of USENIX S...

  13. [20]

    Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Chen Qian, Yujia Qin, Xin Cong, Ruobing Xie, Zhiyuan Liu, Maosong Sun, and Jie Zhou. 2024. AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent B...

  14. [21]

    Gonzalez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. https://lmsys.org/blog/2023-...

  15. [22]

    Yae Jee Cho, Luyang Liu, Zheng Xu, Aldi Fahrezi, and Gauri Joshi. 2024. Heterogeneous LoRA for Federated Fine- tuning of On-Device Foundation Models. In Proceedings of Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguis...

  16. [23]

    Byung-Gon Chun, Sunghwan Ihm, Petros Maniatis, Mayur Naik, and Ashwin Patti. 2011. CloneCloud: elastic execution between mobile device and cloud. In Proceedings of European conference on Computer systems (EuroSys) . ACM, Salzburg, Austria, 301–314

  17. [24]

    Kamil Ciosek. 2022. Imitation Learning by Reinforcement Learning. In Proceedings of International Conference on Learning Representations (ICLR). OpenReview.net, Virtual, 1–15. 24 Niu et al

  18. [25]

    Liam Collins, Hamed Hassani, Aryan Mokhtari, and Sanjay Shakkottai. 2021. Exploiting Shared Representations for Personalized Federated Learning. In Proceedings of International Conference on Machine Learning (ICML) . PMLR, Virtual, 2089–2099

  19. [26]

    Eduardo Cuervo, Aruna Balasubramanian, Dae-ki Cho, Alec Wolman, Stefan Saroiu, Ranveer Chandra, and Paramvir Bahl. 2010. MAUI: making smartphones last longer with code offload. In Proceedings of International Conference on Mobile Systems, Applications, and Services (MobiSys) ....

  20. [27]

    Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

    Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAttention: Fast and Memory- Efficient Exact Attention with IO-Awareness. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., New Orlean...

  21. [28]

    Robert David, Jared Duke, Advait Jain, Vijay Janapa Reddi, Nat Jeffries, Jian Li, Nick Kreeger, Ian Nappier, Meghna Natraj, Tiezhen Wang, Pete Warden, and Rocky Rhodes. 2021. TensorFlow Lite Micro: Embedded Machine Learning on TinyML Systems. InProceedings of Conference on Mac...

  22. [29]

    DeepSeek-AI. 2025. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. CoRR abs/2501.12948 (2025), 22 pages

  23. [30]

    Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., New Orleans, LA, USA, 15 pages

  24. [31]

    Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. QLoRA: Efficient Finetuning of Quantized LLMs. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., New Orleans, LA, USA, 28 pages

  25. [32]

    Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Rühle, Laks V. S. Lakshmanan, and Ahmed Hassan Awadallah. 2024. Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing. In Proceedings of International Conference on Learning Representations (IC...

  26. [33]

    Yucheng Ding, Chaoyue Niu, Fan Wu, Shaojie Tang, Chengfei Lyu, and Guihai Chen. 2022. On-Device Model Fine-Tuning with Label Correction in Recommender Systems. CoRR abs/2211.01163 (2022), 8 pages

  27. [34]

    Yucheng Ding, Chaoyue Niu, Fan Wu, Shaojie Tang, Chengfei Lyu, and Guihai Chen. 2023. DC-CCL: Device- Cloud Collaborative Controlled Learning for Large Vision Models. CoRR abs/2303.10361 (2023), 12 pages. https: //arxiv.org/abs/2303.10361

  28. [35]

    Yucheng Ding, Chaoyue Niu, Fan Wu, Shaojie Tang, Chengfei Lyu, and Guihai Chen. 2024. Enhancing On-Device LLM Inference with Historical Cloud-Based LLM Interactions. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) . ACM, Barcelona, Spain, 597–608

  29. [36]

    Yucheng Ding, Chaoyue Niu, Fan Wu, Shaojie Tang, Chengfei Lyu, Yanghe Feng, and Guihai Chen. 2022. Federated Submodel Optimization for Hot and Cold Data Features. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., N...

  30. [37]

    Yucheng Ding, Yangwenjian Tan, Xiangyu Liu, Chaoyue Niu, Fandong Meng, Jie Zhou, Ning Liu, Fan Wu, and Guihai Chen. 2025. Personalized Language Model Learning on Text Data Without User Identifiers. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and Data Mining ...

  31. [38]

    Mahoney, and Kurt Keutzer

    Zhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. 2019. HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-Precision. In Proceedings of IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, Seoul, Korea (South), 293–302

  32. [39]

    Kuntai Du, Ahsan Pervaiz, Xin Yuan, Aakanksha Chowdhery, Qizheng Zhang, Henry Hoffmann, and Junchen Jiang

  33. [40]

    Ozdaglar

    Alireza Fallah, Aryan Mokhtari, and Asuman E. Ozdaglar. 2020. Personalized Federated Learning with Theoretical Guarantees: A Model-Agnostic Meta-Learning Approach. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., ...

  34. [41]

    Tiantian Feng, Digbalay Bose, Tuo Zhang, Rajat Hebbar, Anil Ramakrishna, Rahul Gupta, Mi Zhang, Salman Aves- timehr, and Shrikanth Narayanan. 2023. FedMultimodal: A Benchmark for Multimodal Federated Learning. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and ...

  35. [42]

    Raphael Fischer and Amal Saadallah. 2024. AutoXPCR: Automated Multi-Objective Model Selection for Time Series Forecasting. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) . ACM, Barcelona, Spain, 806–815

  36. [43]

    Jonathan Frankle and Michael Carbin. 2019. The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, New Orleans, LA, USA, 42 pages. Collaborative Learning of On-Dev...

  37. [44]

    Elias Frantar and Dan Alistarh. 2023. SparseGPT: Massive Language Models Can be Accurately Pruned in One-Shot. In Proceedings of International Conference on Machine Learning (ICML) , Vol. 202. PMLR, Honolulu, Hawaii, USA, 10323–10337

  38. [45]

    Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. CoRR abs/2210.17323 (2022), 16 pages

  39. [46]

    Jingsheng Gao, Yixin Lian, Ziyi Zhou, Yuzhuo Fu, and Baoyuan Wang. 2023. LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming. InProceedings of Annual Meeting of the Association for Computational Linguistics (ACL). Association for...

  40. [47]

    Georgi Gerganov. 2023. llama.cpp. https://github.com/ggerganov/llama.cpp

  41. [48]

    Chen Gong, Zhenzhe Zheng, Fan Wu, Xiaofeng Jia, and Guihai Chen. 2024. Delta: A Cloud-assisted Data Enrichment Framework for On-Device Continual Learning. InProceedings of Annual International Conference on Mobile Computing and Networking (MobiCom). ACM, Washington D.C., DC, U...

  42. [49]

    Xudong Gong, Qinlin Feng, Yuan Zhang, Jiangling Qin, Weijie Ding, Biao Li, Peng Jiang, and Kun Gai. 2022. Real-time Short Video Recommendation on Mobile Devices. In Proceedings of ACM International Conference on Information & Knowledge Management (CIKM). ACM, Atlanta, GA, USA,...

  43. [50]

    Yu Gong, Ziwen Jiang, Yufei Feng, Binbin Hu, Kaiqi Zhao, Qingwen Liu, and Wenwu Ou. 2020. EdgeRec: Recommender System on Edge in Mobile Taobao. In Proceedings of ACM International Conference on Information and Knowledge Management (CIKM). ACM, Virtual, 2477–2484

  44. [51]

    Google. 2017. TensorFlow Lite or LiteRT. https://ai.google.dev/edge/litert

  45. [52]

    Google. 2024. MediaPipe LLM Inference API. https://ai.google.dev/edge/mediapipe/solutions/genai/llm_inference

  46. [53]

    Gordon, David Ke Hong, Peter M

    Mark S. Gordon, David Ke Hong, Peter M. Chen, Jason Flinn, Scott A. Mahlke, and Zhuoqing Morley Mao. 2015. Accelerating Mobile Applications through Flip-Flop Replication. In Proceedings of Annual International Conference on Mobile Systems, Applications, and Services (MobiSys) ...

  47. [54]

    Gordon, Davoud Anoushe Jamshidi, Scott A

    Mark S. Gordon, Davoud Anoushe Jamshidi, Scott A. Mahlke, Zhuoqing Morley Mao, and Xu Chen. 2012. COMET: Code Offload by Migrating Execution Transparently. In Proceedings of USENIX Symposium on Operating Systems Design and Implementation (OSDI) . USENIX Association, Hollywood,...

  48. [55]

    Renjie Gu, Chaoyue Niu, Yikai Yan, Fan Wu, Shaojie Tang, Rongfeng Jia, Chengfei Lv, and Guihai Chen. 2022. On- Device Learning with Cloud-Coordinated Data Augmentation for Extreme Model Personalization in Recommender Systems. CoRR abs/2201.10382 (2022), 14 pages. https://arxiv...

  49. [56]

    Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2024. MiniLLM: Knowledge Distillation of Large Language Models. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Vienna, Austria, 1–24

  50. [57]

    Yuchao Gu, Xintao Wang, Jay Zhangjie Wu, Yujun Shi, Yunpeng Chen, Zihan Fan, Wuyou Xiao, Rui Zhao, Shuning Chang, Weijia Wu, Yixiao Ge, Ying Shan, and Mike Zheng Shou. 2023. Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion Models. In ...

  51. [58]

    Otkrist Gupta and Ramesh Raskar. 2018. Distributed learning of deep neural network over multiple agents. Journal of Network and Computer Applications 116 (2018), 1–8

  52. [59]

    Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan. 2015. Deep Learning with Limited Numerical Precision. In Proceedings of the 32nd International Conference on Machine Learning (ICML) . JMLR.org, Lille, France, 1737–1746

  53. [60]

    Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. Retrieval Augmented Language Model Pre-Training. InProceedings of International Conference on Machine Learning (ICML). PMLR, Virtual, 3929–3938

  54. [61]

    Song Han, Huizi Mao, and William J. Dally. 2016. Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding. In Proceedings of International Conference on Learning Representations, (ICLR). arXiv, San Juan, Puerto Rico, 14 pages

  55. [62]

    Song Han, Jeff Pool, John Tran, and William J. Dally. 2015. Learning both Weights and Connections for Efficient Neural Networks. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Montreal, Quebec, Canada, 1135–1143

  56. [63]

    Seungyeop Han, Haichen Shen, Matthai Philipose, Sharad Agarwal, Alec Wolman, and Arvind Krishnamurthy. 2016. MCDNN: An Approximation-Based Execution Framework for Deep Stream Processing Under Resource Constraints. In Proceedings of Annual International Conference on Mobile Sys...

  57. [64]

    Hinton, Oriol Vinyals, and Jeffrey Dean

    Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015. Distilling the Knowledge in a Neural Network. CoRR abs/1503.02531 (2015), 9 pages. http://arxiv.org/abs/1503.02531

  58. [65]

    Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-Efficient Transfer Learning for NLP. InProceedings of International Conference on Machine Learning (ICML) . PMLR, Lon...

  59. [66]

    Le, Mark Sandler, Bo Chen, Weijun Wang, Liang-Chieh Chen, Mingxing Tan, Grace Chu, Vijay Vasudevan, and Yukun Zhu

    Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le, Mark Sandler, Bo Chen, Weijun Wang, Liang-Chieh Chen, Mingxing Tan, Grace Chu, Vijay Vasudevan, and Yukun Zhu. 2019. Searching for MobileNetV3. In Proceedings of IEEE/CVF International Conference on Computer Vision (ICCV) ...

  60. [67]

    Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam

    Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. CoRR abs/1704.04861 (2017), 9 pages. https://arxiv.org/abs/1704.04861

  61. [68]

    Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. 2023. Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes. In Findings of the As...

  62. [69]

    Tzu-Ming Harry Hsu, Hang Qi, and Matthew Brown. 2020. Federated Visual Classification with Real-World Data Distribution. In Proceedings of European Conference on Computer Vision (ECCV) . Springer, Glasgow, UK, 76–92

  63. [70]

    Bo Hu and Wenjun Hu. 2019. LinkShare: device-centric control for concurrent and continuous mobile-cloud interactions. In Proceedings of ACM/IEEE Symposium on Edge Computing (SEC) . ACM, Arlington, Virginia, USA, 15–29

  64. [71]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

  65. [72]

    Kai Huang and Wei Gao. 2022. Real-time neural network inference on extremely weak devices: agile offloading with explainable AI. In Proceedings of Annual International Conference on Mobile Computing and Networking (MobiCom) . ACM, Sydney, NSW, Australia, 200–213

  66. [73]

    Wenke Huang, Mang Ye, Zekun Shi, He Li, and Bo Du. 2023. Rethinking Federated Learning with Domain Shift: A Prototype View. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, Vancouver, BC, Canada, 16312–16322

  67. [74]

    Hugging Face. 2018. Transformers. https://github.com/huggingface/transformers

  68. [76]

    Howard, Hartwig Adam, and Dmitry Kalenichenko

    Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew G. Howard, Hartwig Adam, and Dmitry Kalenichenko. 2018. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. In Proceedings of IEEE Conference on Computer Visi...

  69. [77]

    Chen Jia. 2024. Adversarial Moment-Matching Distillation of Large Language Models. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Vancouver, BC, Canada, 1–33

  70. [78]

    Bowen Jin, Hansi Zeng, Zhenrui Yue, Dong Wang, Hamed Zamani, and Jiawei Han. 2025. Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning. CoRR abs/2503.09516 (2025), 16 pages

  71. [79]

    Whatmough, and Venkatesh Saligrama

    Anil Kag, Igor Fedorov, Aditya Gangrade, Paul N. Whatmough, and Venkatesh Saligrama. 2023. Efficient Edge Inference by Selective Query. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Kigali, Rwanda, 25 pages

  72. [80]

    Mudge, Jason Mars, and Lingjia Tang

    Yiping Kang, Johann Hauswald, Cao Gao, Austin Rovinski, Trevor N. Mudge, Jason Mars, and Lingjia Tang. 2017. Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge. In Proceedings of International Conference on Architectural Support for Programming Language...

  73. [81]

    Reddi, Sebastian U

    Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi, Sebastian U. Stich, and Ananda Theertha Suresh. 2020. SCAFFOLD: Stochastic Controlled Averaging for Federated Learning. In Proceedings of International Conference on Machine Learning (ICML) , Vol. 119. PML...

  74. [82]

    Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020. Generalization through Mem- orization: Nearest Neighbor Language Models. In Proceedings of International Conference on Learning Representations (ICLR). OpenReview.net, Addis Ababa, Ethiopia, 13 pages

  75. [83]

    Yoon Kim and Alexander M. Rush. 2016. Sequence-Level Knowledge Distillation. In Proceedings of Conference on Empirical Methods in Natural Language Processing (EMNLP) . The Association for Computational Linguistics, Austin, Texas, USA, 1317–1327

  76. [84]

    Jongwoo Ko, Sungnyun Kim, Tianyi Chen, and Se-Young Yun. 2024. DistiLLM: Towards Streamlined Distillation for Large Language Models. In Proceedings of International Conference on Machine Learning (ICML) . OpenReview.net, Vienna, Austria, 1–24

  77. [85]

    Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian U. Stich. 2020. A Unified Theory of Decentralized SGD with Changing Topology and Local Updates. In Proceedings of International Conference on Machine Learning (ICML). PMLR, Virtual, 5381–5393. Coll...

  78. [86]

    Venieris, Mário Almeida, Ilias Leontiadis, and Nicholas D

    Stefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis, and Nicholas D. Lane. 2020. SPINN: synergistic progressive inference of neural networks over device and cloud. In Proceedings of Annual International Conference on Mobile Computing and Networking (Mob...

  79. [87]

    Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The Power of Scale for Parameter-Efficient Prompt Tun- ing. In Proceedings of Conference on Empirical Methods in Natural Language Processing (EMNLP) . Association for Computational Linguistics, Punta Cana, Dominican Republic...

  80. [88]

    Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast Inference from Transformers via Speculative Decoding. In International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 202) . Proceedings of Machine Learning Research (PMLR), Hono...

  81. [89]

    Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. InProceed...

  82. [90]

    Ang Li, Jingwei Sun, Pengcheng Li, Yu Pu, Hai Li, and Yiran Chen. 2021. Hermes: an efficient federated learning framework for heterogeneous mobile clients. In Proceedings of Annual International Conference on Mobile Computing and Networking (MobiCom). ACM, New Orleans, Louisia...

  83. [91]

    Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023. CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, I...

  84. [92]

    Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren, Kai-Wei Chang, and Yejin Choi. 2023. Symbolic Chain-of- Thought Distillation: Small Models Can Also "Think" Step-by-Step. InProceedings of Annual Meeting of the Association for Computational Linguistics (ACL). Association...

  85. [93]

    Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith. 2020. Federated Optimization in Heterogeneous Networks. In Proceedings of Conference on Machine Learning and Systems (MLSys) . mlsys.org, Austin, TX, USA, 22 pages

  86. [94]

    Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Zhihua Zhang. 2020. On the Convergence of FedAvg on Non-IID Data. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Addis Ababa, Ethiopia, 26 pages

  87. [95]

    Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2023. Contrastive Decoding: Open-ended Text Generation as Optimization. In Proceedings of Annual Meeting of the Association for Computational Linguistic...

  88. [96]

    Xiang Lisa Li and Percy Liang. 2021. Prefix-Tuning: Optimizing Continuous Prompts for Generation. In Proceedings of Annual Meeting of the Association for Computational Linguistics (ACL) . ACL, Virtual, 4582–4597

  89. [97]

    Yuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Yufei Wang, Guoqing Harry Xu, and Ravi Netravali. 2020. Reducto: On-camera Filtering for Resource-Efficient Real-Time Video Analytics. InProceedings of the ACM Special Interest Group on Data Communication on the applications, techno...

  90. [98]

    Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset. In Proceedings of International Joint Conference on Natural Language Processing (IJCNLP) . Asian Federation of Natural Language Processing...

  91. [99]

    Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2024. AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration. In Proceedings of the Seventh Annual Conference ...

  92. [100]

    Stich, and Martin Jaggi

    Tao Lin, Lingjing Kong, Sebastian U. Stich, and Martin Jaggi. 2020. Ensemble Distillation for Robust Model Fusion in Federated Learning. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., virtual, 13 pages

  93. [101]

    Alisa Liu, Xiaochuang Han, Yizhong Wang, Yulia Tsvetkov, Yejin Choi, and Noah A. Smith. 2024. Tuning Lan- guage Models by Proxy. In Proceedings of Conference on Language Modeling (COLM) . OpenReview.net, Philadelphia, Pennsylvania, USA, 24 pages

  94. [102]

    Jiaheng Liu, Chenchen Zhang, Jinyang Guo, Yuanxing Zhang, Haoran Que, Ken Deng, Zhiqi Bai, Jie Liu, Ge Zhang, Jiakai Wang, Yanan Wu, Congnan Liu, Jiamang Wang, Lin Qu, Wenbo Su, and Bo Zheng. 2024. DDK: Distilling Domain Knowledge for Efficient Large Language Models. In Procee...

  95. [103]

    Yuejiang Liu and Alexandre Alahi. 2024. Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts. CoRR abs/2402.15505 (2024), 9 pages. 28 Niu et al

  96. [104]

    Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. 2017. Learning Efficient Convolutional Networks through Network Slimming. In Proceedings of IEEE International Conference on Computer Vision (ICCV). IEEE Computer Society, Venice, Italy, 2755–2763

  97. [105]

    Zechun Liu, Barlas Oguz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, and Vikas Chandra. 2024. LLM-QAT: Data-Free Quantization Aware Training for Large Language Models. In Findings of the Association for Computational Ling...

  98. [106]

    Llama Team, AI @ Meta. 2024. The Llama 3 Herd of Models. CoRR abs/2407.21783 (2024), 92 pages. https: //arxiv.org/abs/2407.21783

  99. [107]

    Keming Lu, Hongyi Yuan, Runji Lin, Junyang Lin, Zheng Yuan, Chang Zhou, and Jingren Zhou. 2024. Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models. In Proceedings of Conference of the North American Chapter of the Association for Computational Lin...

  100. [108]

    Yan Lu, Yuanchao Shu, Xu Tan, Yunxin Liu, Mengyu Zhou, Qi Chen, and Dan Pei. 2019. Collaborative learning between cloud and end devices: an empirical study on location prediction. InProceedings of the ACM/IEEE Symposium on Edge Computing (SEC) . ACM/IEEE, Washington DC, USA, 139–151

  101. [109]

    Chengfei Lv, Chaoyue Niu, Renjie Gu, Xiaotang Jiang, Zhaode Wang, Bin Liu, Ziqi Wu, Qiulin Yao, Congyu Huang, Panos Huang, Tao Huang, Hui Shu, Jinde Song, Bin Zou, Peng Lan, Guohuan Xu, Fei Wu, Shaojie Tang, Fan Wu, and Guihai Chen. 2022. Walle: An End-to-End, General-Purpose,...

  102. [110]

    Zheqi Lv, Tianyu Zhan, Wenjie Wang, Xinyu Lin, Shengyu Zhang, Wenqiao Zhang, Jiwei Li, Kun Kuang, and Fei Wu. 2025. Collaboration of Large Language Models and Small Recommendation Models for Device-Cloud Recommendation. In Proceedings of ACM SIGKDD Conference on Knowledge Disc...

  103. [111]

    Zheqi Lv, Wenqiao Zhang, Shengyu Zhang, Kun Kuang, Feng Wang, Yongwei Wang, Zhengyu Chen, Tao Shen, Hongxia Yang, Beng Chin Ooi, and Fei Wu. 2023. DUET: A Tuning-Free Device-Cloud Collaborative Parameters Generation Framework for Efficient Device Model Generalization. In Proce...

  104. [112]

    Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2023. LLM-Pruner: On the Structural Pruning of Large Language Models. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., New Orleans, LA, USA, 19 pages

  105. [113]

    Xiaosong Ma, Jie Zhang, Song Guo, and Wenchao Xu. 2022. Layer-wised Model Aggregation for Personalized Federated Learning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, New Orleans, LA, USA, 10082–10091

  106. [114]

    Yishay Mansour, Mehryar Mohri, Jae Ro, and Ananda Theertha Suresh. 2020. Three Approaches for Personalization with Applications to Federated Learning. CoRR abs/2002.10619 (2020), 26 pages

  107. [115]

    Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas. 2017. Communication- Efficient Learning of Deep Networks from Decentralized Data. In Proceedings of International Conference on Artificial Intelligence and Statistics (AISTATS). PMLR, Fort La...

  108. [116]

    Meta. 2019. PyTorch Mobile. https://pytorch.org/mobile/home/

  109. [117]

    Meta. 2023. ExecuTorch. https://github.com/pytorch/executorch

  110. [118]

    Meta. 2024. torchchat. https://github.com/pytorch/torchchat

  111. [119]

    Microsoft. 2020. DeepSpeed. https://github.com/microsoft/DeepSpeed

  112. [120]

    Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius

    Asit K. Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius. 2021. Accelerating Sparse Deep Neural Networks. CoRR abs/2104.08378 (2021), 18 pages

  113. [121]

    Eric Mitchell, Rafael Rafailov, Archit Sharma, Chelsea Finn, and Christopher D. Manning. 2024. An Emulator for Fine-tuning Large Language Models using Small Language Models. In Proceedings of International Conference on Learning Representations (ICLR). OpenReview.net, Vienna, ...

  114. [122]

    Arindam Mitra, Hamed Khanpour, Corby Rosset, and Ahmed Awadallah. 2024. Orca-Math: Unlocking the potential of SLMs in Grade School Math. CoRR abs/2402.14830 (2024), 14 pages

  115. [123]

    Chaoyue Niu, Fan Wu, Shaojie Tang, Lifeng Hua, Rongfei Jia, Chengfei Lv, Zhihua Wu, and Guihai Chen. 2020. Billion-scale federated learning on mobile clients: a submodel design with tunable privacy. In Proceedings of Annual International Conference on Mobile Computing and Netw...

  116. [124]

    NVIDIA. 2019. Megatron-LM. https://github.com/NVIDIA/Megatron-LM

  117. [125]

    Gonzalez, M Waleed Kadous, and Ion Stoica

    Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, and Ion Stoica. 2025. RouteLLM: Learning to Route LLMs from Preference Data. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net,...

  118. [126]

    Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, ...

  119. [127]

    Ryan Po, Guandao Yang, Kfir Aberman, and Gordon Wetzstein. 2024. Orthogonal Adaptation for Modular Customiza- tion of Diffusion Models. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, Seattle, WA, USA, 7964–7973

  120. [128]

    Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2024. ChatDev: Communicative Agents for Software Development. In Proceedings of Annual Meeting of the Associatio...

  121. [129]

    Qualcomm. 2024. The future of AI is hybrid; Part I: Unlocking the generative AI future with on-device and hybrid AI. https://www.qualcomm.com/content/dam/qcomm-martech/dm-assets/documents/Whitepaper-The-future-of- AI-is-hybrid-Part-1-Unlocking-the-generative-AI-future-with-on-...

  122. [130]

    Qualcomm. 2024. The future of AI is hybrid; Part II: Qualcomm is uniquely positioned to scale hybrid AI. https://www.qualcomm.com/content/dam/qcomm-martech/dm-assets/documents/Whitepaper-The-future-of- AI-is-hybrid-Part-2-Qualcomm-is-uniquely-positioned-to-scale-hybrid-AI.pdf

  123. [131]

    Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. 2016. XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks. In Proceedings of European Conference on Computer Vision (ECCV) . Springer, Amsterdam, The Netherlands, 525–542

  124. [132]

    Reddi, Aditya Krishna Menon, Rohan Anil, and Sanjiv Kumar

    Ankit Singh Rawat, Veeranjaneyulu Sadhanala, Afshin Rostamizadeh, Ayan Chakrabarti, Wittawat Jitkrittum, Vladimir Feinberg, Seungyeon Kim, Hrayr Harutyunyan, Nikunj Saunshi, Zachary Nado, Rakesh Shivanna, Sashank J. Reddi, Aditya Krishna Menon, Rohan Anil, and Sanjiv Kumar. 20...

  125. [133]

    Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2015. FitNets: Hints for Thin Deep Nets. CoRR abs/1412.6550 (2015), 13 pages

  126. [134]

    Yichen Ruan, Xiaoxi Zhang, Shu-Che Liang, and Carlee Joe-Wong. 2021. Towards Flexible Device Participation in Federated Learning. In Proceedings of International Conference on Artificial Intelligence and Statistics (AISTATS) . PMLR, Virtual, 3403–3411

  127. [135]

    Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen

    Mark Sandler, Andrew G. Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. 2018. MobileNetV2: Inverted Residuals and Linear Bottlenecks. InProceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Computer Vision Foundation / IEEE Computer So...

  128. [136]

    Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Tran, Yi Tay, and Donald Metzler. 2022. Confident Adaptive Language Modeling. InProceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., New Orleans, LA...

  129. [137]

    Virginia Smith, Chao-Kai Chiang, Maziar Sanjabi, and Ameet Talwalkar. 2017. Federated Multi-Task Learning. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Long Beach, CA, USA, 4424–4434

  130. [138]

    Nimit Sharad Sohoni, Christopher Richard Aberger, Megan Leszczynski, Jian Zhang, and Christopher Ré. 2019. Low-Memory Neural Network Training: A Technical Report. CoRR abs/1904.10631 (2019), 38 pages

  131. [139]

    Nikita Starodubcev, Dmitry Baranchuk, Artem Fedorov, and Artem Babenko. 2024. Your Student is Better than Expected: Adaptive Teacher-Student Collaboration for Text-Conditional Diffusion Models. InProceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CV...

  132. [140]

    Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V. Le. 2019. MnasNet: Platform-Aware Neural Architecture Search for Mobile. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Computer Vision Foundat...

  133. [141]

    Mingxing Tan and Quoc V. Le. 2019. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of International Conference on Machine Learning (ICML) . PMLR, Long Beach, CA, USA, 6105–6114

  134. [142]

    Yue Tan, Guodong Long, Lu Liu, Tianyi Zhou, Qinghua Lu, Jing Jiang, and Chengqi Zhang. 2022. FedProto: Federated Prototype Learning across Heterogeneous Clients. In Proceedings of AAAI Conference on Artificial Intelligence (AAAI) . AAAI Press, Virtual, 8432–8440

  135. [143]

    Yue Tan, Guodong Long, Jie Ma, Lu Liu, Tianyi Zhou, and Jing Jiang. 2022. Federated Learning from Pre-Trained Models: A Contrastive Learning Approach. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., New Orleans, ...

  136. [144]

    Hashimoto

    Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford Alpaca: An Instruction-following LLaMA model. https://github.com/tatsu-lab/stanford_ 30 Niu et al. alpaca

  137. [145]

    MLC team. 2023. MLC-LLM. https://github.com/mlc-ai/mlc-llm

  138. [146]

    Surat Teerapittayanon, Bradley McDanel, and H. T. Kung. 2016. BranchyNet: Fast inference via early exiting from deep neural networks. In Proceedings of International Conference on Pattern Recognition (ICPR) . IEEE, Cancún, Mexico, 2464–2469

  139. [147]

    Surat Teerapittayanon, Bradley McDanel, and H. T. Kung. 2017. Distributed Deep Neural Networks Over the Cloud, the Edge and End Devices. In Proceedings of IEEE International Conference on Distributed Computing Systems (ICDCS) . IEEE, Atlanta, GA, USA, 328–339

  140. [148]

    Vincent Poor

    Jianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi, and H. Vincent Poor. 2020. Tackling the Objective Inconsistency Problem in Heterogeneous Federated Optimization. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., V...

  141. [149]

    Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han. 2019. HAQ: Hardware-Aware Automated Quantization With Mixed Precision. InProceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Computer Vision Foundation / IEEE, Long Beach, CA, USA, 8612–8620

  142. [150]

    Qipeng Wang, Mengwei Xu, Chao Jin, Xinran Dong, Jinliang Yuan, Xin Jin, Gang Huang, Yunxin Liu, and Xuanzhe Liu. 2022. Melon: breaking the memory wall for resource-efficient on-device machine learning. In Proceedings of Annual International Conference on Mobile Systems, Applic...

  143. [151]

    Yiding Wang, Kai Chen, Haisheng Tan, and Kun Guo. 2023. Tabi: An Efficient Multi-Level Inference System for Large Language Models. In Proceedings of European Conference on Computer Systems (EuroSys) . ACM, Rome, Italy, 233–248

  144. [152]

    Heming Xia, Tao Ge, Peiyi Wang, Si-Qing Chen, Furu Wei, and Zhifang Sui. 2023. Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation. InFindings of the Association for Computational Linguistics (EMNLP). Association for Computational Linguis...

  145. [153]

    Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen. 2024. Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning. InProceedings of The Twelfth International Conference on Learning Representations (ICLR). OpenReview.net, Vienna, Austria, 25 pages

  146. [154]

    Guangxuan Xiao, Ji Lin, and Song Han. 2023. Offsite-Tuning: Transfer Learning without Full Model. CoRR abs/2302.04870 (2023), 12 pages. https://arxiv.org/abs/2302.04870

  147. [155]

    Guangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu, Julien Demouth, and Song Han. 2023. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models. In Proceedings of International Conference on Machine Learning (ICML), Vol. 202. PMLR, Honolulu, Hawaii...

  148. [156]

    Hovy, and Quoc V

    Qizhe Xie, Minh-Thang Luong, Eduard H. Hovy, and Quoc V. Le. 2020. Self-Training With Noisy Student Improves ImageNet Classification. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Computer Vision Foundation / IEEE, Seattle, WA, USA, ...

  149. [157]

    Zewei Xin, Qinya Li, Chaoyue Niu, and Fan Wu. 2024. Edge-Cloud Routing for Text-to-Image Model with Token-Level Multi-Metric Prediction. CoRR abs/2411.13787 (2024), 10 pages. https://doi.org/10.48550/arXiv.2411.13787

  150. [158]

    Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang

  151. [159]

    Xinchao Xu, Zhibin Gou, Wenquan Wu, Zheng-Yu Niu, Hua Wu, Haifeng Wang, and Shihang Wang. 2022. Long Time No See! Open-Domain Conversation with Long-Term Persona Memory. In Findings of the Association for Computational Linguistics (ACL). Association for Computational Linguisti...

  152. [160]

    Yikai Yan, Chaoyue Niu, Yucheng Ding, Zhenzhe Zheng, Shaojie Tang, Qinya Li, Fan Wu, Chengfei Lyu, Yanghe Feng, and Guihai Chen. 2024. Federated Optimization Under Intermittent Client Availability. INFORMS Journal on Computing 36, 1 (2024), 185–202

  153. [161]

    Yikai Yan, Chaoyue Niu, Renjie Gu, Fan Wu, Shaojie Tang, Lifeng Hua, Chengfei Lyu, and Guihai Chen. 2022. On- Device Learning for Model Personalization with Large-Scale Cloud-Coordinated Domain Adaption. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and Data M...

  154. [162]

    Jiangchao Yao, Feng Wang, Kunyang Jia, Bo Han, Jingren Zhou, and Hongxia Yang. 2021. Device-Cloud Collaborative Learning for Recommendation. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD). ACM, Virtual, 3865–3874

  155. [163]

    Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. 2022. ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers. InProceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) ....

  156. [164]

    In Proceedings of International Conference on Learning Representations (ICLR)

    WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex Instructions. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Vienna, Austria, 1–22

  157. [165]

    Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu

    Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2024. MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models. In Proceedings of International Conference on Learning Repre...

  158. [166]

    Murong Yue, Jie Zhao, Min Zhang, Liang Du, and Ziyu Yao. 2024. Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient Reasoning. In Proceedings of International Conference on Learning Representations (ICLR). OpenReview.net, Vienna, Austria, 38 pages

  159. [167]

    Kaiyan Zhang, Jianyu Wang, Ning Ding, Biqing Qi, Ermo Hua, Xingtai Lv, and Bowen Zhou. 2024. Fast and Slow Generating: An Empirical Study on Large and Small Language Models Collaborative Decoding. CoRR abs/2406.12295 (2024), 17 pages. https://arxiv.org/abs/2406.12295

  160. [168]

    Kaiyan Zhang, Jianyu Wang, Ermo Hua, Biqing Qi, Ning Ding, and Bowen Zhou. 2024. CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction Following. In Proceedings of Annual Meeting of the Association for Computational Linguisti...

  161. [169]

    Michael Zhang, Karan Sapra, Sanja Fidler, Serena Yeung, and José M. Álvarez. 2021. Personalized Federated Learning with First Order Model Optimization. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Virtual, 17 pages

  162. [170]

    Junho Yim, Donggyu Joo, Ji-Hoon Bae, and Junmo Kim. 2017. A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer Learning. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE Computer Society, Honolulu, HI...

  163. [171]

    Hospedales, and Huchuan Lu

    Ying Zhang, Tao Xiang, Timothy M. Hospedales, and Huchuan Lu. 2018. Deep Mutual Learning. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Computer Vision Foundation / IEEE Computer Society, Salt Lake City, UT, USA, 4320–4328

  164. [172]

    Yinhe Zheng, Guanyi Chen, Minlie Huang, Song Liu, and Xuan Zhu. 2019. Personalized Dialogue Generation with Diversified Traits. CoRR abs/1901.09672 (2019), 12 pages

  165. [173]

    Qihuang Zhong, Liang Ding, Li Shen, Juhua Liu, Bo Du, and Dacheng Tao. 2024. Revisiting Knowledge Distillation for Autoregressive Language Models. In Proceedings of Annual Meeting of the Association for Computational Linguistics (ACL). Association for Computational Linguistics...

  166. [174]

    Shuchang Zhou, Zekun Ni, Xinyu Zhou, He Wen, Yuxin Wu, and Yuheng Zou. 2016. DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients. CoRR abs/1606.06160 (2016), 13 pages

  167. [175]

    Zhanhui Zhou, Zhixuan Liu, Jie Liu, Zhichen Dong, Chao Yang, and Yu Qiao. 2024. Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associat...

  168. [176]

    Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Personalizing Dialogue Agents: I have a dog, do you have pets too?. In Proceedings of Annual Meeting of the Association for Computational Linguistics (ACL). Association for Computatio...

  169. [177]

    Han Zhu, Junqi Jin, Chang Tan, Fei Pan, Yifan Zeng, Han Li, and Kun Gai. 2017. Optimized Cost per Click in Taobao Display Advertising. In Proceedings of ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD). ACM, Halifax, NS, Canada, 2191–2200

  170. [178]

    Yufei Zhu, Chaoyue Niu, Yikai Yan, Zhijie Cao, Hao Jiang, Chengfei Lyu, Shaojie Tang, and Fan Wu. 2023. Device- Unimodal Cloud-Multimodal Collaboration for Livestreaming Content Understanding. In Proceedings of IEEE Interna- tional Conference on Data Mining (ICDM) . IEEE, Shan...

  171. [179]

    Zhuangdi Zhu, Junyuan Hong, and Jiayu Zhou. 2021. Data-Free Knowledge Distillation for Heterogeneous Federated Learning. In Proceedings of International Conference on Machine Learning (ICML) . PMLR, Virtual, 12878–12889

  172. [180]

    Richard Zhuang, Tianhao Wu, Zhaojin Wen, Andrew Li, Jiantao Jiao, and Kannan Ramchandran. 2025. EmbedLLM: Learning Compact Representations of Large Language Models. In Proceedings of International Conference on Learning Representations (ICLR). OpenReview.net, Singapore, 14 pages

  173. [181]

    Barret Zoph and Quoc V. Le. 2017. Neural Architecture Search with Reinforcement Learning. In Proceedings of International Conference on Learning Representations, (ICLR) . OpenReview.net, Toulon, France, 16 pages

  174. [182]

    Barrett, Michael I

    Banghua Zhu, Ying Sheng, Lianmin Zheng, Clark W. Barrett, Michael I. Jordan, and Jiantao Jiao. 2023. On Optimal Caching and Model Selection for Large Model Inference. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc...

  175. [2016]

    In Proceedings of USENIX Symposium on Operating Systems Design and Implementation (OSDI)

    TensorFlow: A System for Large-Scale Machine Learning. In Proceedings of USENIX Symposium on Operating Systems Design and Implementation (OSDI) . USENIX, Savannah, GA, USA, 265–283

  176. [2018]

    CoRR abs/1812.01097 (2018), 9 pages

    LEAF: A Benchmark for Federated Settings. CoRR abs/1812.01097 (2018), 9 pages

  177. [2020]

    In Proceedings of Annual Conference of ACM’s Special Interest Group on Data Communication (SIGCOMM)

    Server-driven video streaming for deep learning inference. In Proceedings of Annual Conference of ACM’s Special Interest Group on Data Communication (SIGCOMM) . ACM, Virtual, 557–570

  178. [2022]

    In Proceedings of International Conference on Learning Representations (ICLR)

    LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of International Conference on Learning Representations (ICLR). OpenReview.net, Virtual, 13 pages

  179. [2024]

    In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS)

    AutoMix: Automatically Mixing Language Models. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS). Curran Associates, Inc., Vancouver, BC, Canada, 35 pages

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.