REVIEW 3 major objections 5 minor 162 references
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read CLONE is an algorithm-hardware co-design that speeds edge LLM inference up to 11.92x and cuts energy up to 7.36x while preserving generation quality, combining generative structural pruning, prompt-based LoRA routing, layer-wise DVFS, and…
desk verdict Competent integration of known edge-LLM techniques whose headline numbers rest on a simulated accelerator and a strawman baseline; the measured software gains are modest and the paper deserves conditional review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the ratio-score metric $s_i=f(r_i)$ from Eq. (1), which scores a candidate layer-wise pruning ratio $r_i$ by generation quality (inverse perplexity) multiplied by budget terms that penalize exceeding latency or energy targets; this score turns pruning into an optimizable continuous-space search. The search is carried by an encoder-evaluator-decoder built from a single-layer LSTM encoder/decoder and feed-forward evaluator, which embeds collected ratio-score pairs into a continuous space and follows gradients from top-K starting points to generate the optimal pruning configuration with beam search. Online, the system's second mechanism is a parameter-free Mixture-of-Experts (MoE) router: a sentence-embedding model computes cosine similarity between the request and each LoRA adapter, then a softmax mixes adapter outputs without trainable gates. The third mechanism is a learning-based, layer-wise dynamic voltage and frequency scaling (DVFS) controller, a two-layer MLP trained by reinforcement learning, which chooses voltage/frequency per layer per token to meet latency targets while minimizing energy. These mechanisms are mapped onto a 28nm accelerator with a LoRA Processing Unit (LPU) for adapter hot-swapping and a Special Function Unit (SFU) with a fast-switching low-dropout (LDO) voltage regulator and an all-digital phase-locked loop (ADPLL) for fine-grained DVFS.
What would settle it
Fabricate or FPGA-emulate the 28nm accelerator, measure the LoRA processing unit's throughput and the LDO/ADPLL switching times, and recompute end-to-end latency and energy on the Jetson platforms; if measured switching times or throughput differ from the post-layout values, the headline speedup and energy savings change by the same margin.
Extended reading notes
Core claim
CLONE's central claim is that LLM inference at the edge should be treated as a joint model-system-hardware optimization instead of separate compression and scheduling steps. The paper shows that decoder layers contribute unevenly to generation quality, latency, and energy, and it exploits this by reframing pruning as a generative task: an encoder-evaluator-decoder learns a continuous space of pruning-ratio configurations using a holistic score $s_i=f(r_i)$ that combines zero-shot perplexity with latency and energy budgets, then gradient-based search produces the layer-wise pruning configuration. Online, a parameter-free Mixture-of-Experts router uses cosine similarity between a sentence embedding of the request and embeddings of LoRA adapters to select and mix adapters per request, while a reinforcement-learning controller applies dynamic voltage and frequency scaling at layer boundaries for each generated token. A 28nm accelerator with a LoRA Processing Unit and a Special Function Unit implements the router and fine-grained DVFS. On two off-the-shelf Jetson platforms, the end-to-end design accelerates inference by up to 11.92x and reduces energy by up to 7.36x while keeping generation quality close to the unpruned model.
Load-bearing premise
The headline 11.92x and 7.36x figures assume that post-layout simulation of the 28nm accelerator predicts real silicon behavior, because the chip was validated only in simulation and never fabricated.
Editorial extensions
If this is right
- Jetson-class devices could serve interactive LLM workloads that previously required cloud offloading, bringing end-to-end latency closer to the human-acceptable thresholds the paper cites.
- A single pruned base model can cover many applications by swapping and mixing LoRA adapters per request, avoiding full-model retraining for each new task.
- Layer-wise DVFS during per-token decoding decouples prefill and decode energy control, enabling finer-grained trade-offs than whole-model black-box scaling.
- Without the accelerator, CLONE-HW still beats all baselines on energy and latency (4.81 Wh and 462.72 s on Orin NX), so the hardware units add a measurable share beyond the algorithmic gains.
Reading between the lines
- Editorial inference: the parameter-free MoE routing rule could adapt to new or drifting tasks without retraining the router, since it depends only on embedding similarity, but the paper does not test this.
- Editorial inference: the 11.92x and 7.36x figures mix measured Jetson results with post-layout simulated accelerator results, so on fabricated silicon the split between algorithmic and hardware gains may shift; the headline should be read as a simulation-anchored estimate until the chip is taped out.
- Editorial inference: the ratio-score metric suggests a testable coupling between offline pruning and online adaptation, updating pruning ratios as request distributions drift; the paper keeps the two phases separate.
- Editorial inference: a head-to-head comparison against 4-bit quantized baselines under the same latency targets would isolate how much of CLONE's gain comes from pruning plus DVFS rather than quantization, a baseline family the paper does not include.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents CLONE, an algorithm-hardware co-design system for edge LLM inference. The software side combines offline generative pruning (an encoder-evaluator-decoder that searches layer-wise pruning ratios against an objective including perplexity, latency, and energy), plug-and-play LoRA adapters with a request-wise MoE router, and a learning-based layer-wise DVFS controller. The hardware side is a 28nm accelerator containing a LoRA Processing Unit and a Special Function Unit, validated through post-layout simulation rather than fabrication. The authors evaluate on Jetson Orin NX and Orin Nano with Llama-7B, Llama2-7B, Llama2-13B, and Vicuna-7B, reporting latency and energy improvements up to 11.92x and 7.36x respectively, while preserving or improving downstream benchmark accuracy relative to pruning baselines.
Significance. If the claims hold, CLONE would be a significant demonstration of model-system-hardware co-design for edge LLM deployment, with a useful combination of pruning, multi-task adapters, and DVFS. The paper has real strengths: it is evaluated on two commercial Jetson devices with multiple LLMs; downstream quality is assessed on independent benchmarks (BBH, MMLU, commonsense reasoning); the soft MoE router is parameter-free and overhead is quantified; and the offline tailoring hyperparameters are reported in reasonable detail. The main risk is that the headline speedup and energy numbers are not reproducible from the evidence as presented, because they rely on a simulated accelerator mixed with measured Jetson execution, and because the reference baseline (FlexGen) is the slowest in the comparison table. The measured software-only contribution (CLONE-HW) is more modest but still positive against pruning baselines, which suggests the algorithmic core is defensible and can be rescued by a more careful reporting of simulated versus measured results.
major comments (3)
- [Section 5.2, Table 3, Abstract] The headline 'up to 11.92x/7.36x' figures in the abstract and Section 5.3 refer to the CLONE row of Table 3, but Section 5.2 states that the 28nm accelerator was only 'validated through post-layout simulation.' The paper never explains how simulated LDO/ADPLL switching times and LPU throughput were combined with measured Jetson execution to produce the CLONE row (322.76 s / 3.46 Wh on Orin NX; 392.15 s / 3.54 Wh on Orin Nano). As written, these central numbers mix measured and simulated quantities without a stated methodology, so they cannot be independently reproduced. The authors should present measured software-only results as the primary evidence and clearly label the simulated hardware contribution as a projection, with the integration methodology and simulation assumptions specified.
- [Table 3, Section 5.3, Abstract] The reported speedup is anchored to FlexGen, which in this evaluation keeps full-model weights in CPU DRAM and reloads them per layer, making it an unusually slow baseline for edge inference. The same table shows that CLONE-HW is only about 1.2x-1.8x faster than LLMPruner, ShortGPT, or SliceGPT on latency, and the energy advantage over those baselines is similarly modest. Since the abstract and conclusion emphasize the 11.92x and 7.36x figures without stating that these are FlexGen-relative and include simulated hardware, the paper materially overstates the practical improvement. The authors should headline comparisons against the stronger pruning baselines or report both with explicit conditioning on the baseline and the measurement basis.
- [Equation (1), Table 3] The pruning objective in Eq. (1) directly includes inference latency and energy, so the later system-effectiveness comparison in Table 3 is partly circular: CLONE is optimized to improve exactly the metrics on which it is then evaluated. The independent downstream benchmarks (BBH, MMLU, commonsense) do not have this issue, and the model-quality claims survive it. However, the latency/energy claims need an ablation or a validation procedure that separates the effect of the optimization objective from the effect of the pruning configuration itself, for example by evaluating the generative search against same-budget pruned models that do not incorporate latency and energy terms in their scoring.
minor comments (5)
- [Abstract, Section 5.3] The phrase 'maintaining high-generation' is incomplete; it should read 'maintaining high generation quality.'
- [Table 3 caption] The caption does not explain the difference between CLONE and CLONE-HW; 'CLONE-HW' is defined only in the body text, making the table hard to interpret in isolation.
- [Figure 17] The x-axis labels of Figure 17 are rendered as unreadable glyphs; they should be replaced with clean text labels so the layer-wise pruning ratios can actually be read.
- [References] References [74] and [75] are duplicate entries for the same PagedAttention paper and should be merged.
- [Section 4.4] The hardware section describes the accelerator as supporting 'hot-swapping' and fast DVFS, but since the chip was not fabricated, the text should consistently state that these are simulated capabilities rather than measured silicon behavior.
Circularity Check
Only partial circularity: the pruning objective in Eq. (1) includes the same perplexity/latency/energy metrics later reported as success; independent downstream benchmarks prevent the central claim from being forced.
-
fitted input called prediction
[Section 4.2, Eq. (1) and Section 5.3, 'System Effectiveness' and Table 3]
"si =f(r i) = 1 ppli × (E ei )1(E<ei)×α × (T ti )1(T<ti)×β ... As shown in Equation 1, CLONE jointly considers the latency and energy budgets as metrics to generate optimal pruning configuration."
Eq. (1) defines the tailoring score using zero-shot perplexity, measured latency ti, and measured energy ei, and the generative tailor is trained on ratio-score pairs and optimized over this score. The System Effectiveness section then reports WikiText2 latency and energy (Table 3) as evidence of 'up to 11.92x' speedup and 'up to 7.36x' energy saving, and the Generation Ability section reports WikiText2/PTB PPL. These are the same quantities that were optimized in Eq. (1), so the reported gains in those metrics are partly induced by the selection criterion rather than independent predictions. The central claim retains independent content because BBH, MMLU, and Commonsense benchmarks are not part of Eq.
full rationale
The paper's only significant self-referential element is that the pruning objective in Eq. (1) is built from the same perplexity, latency, and energy metrics that the evaluation then reports as gains. This is a real but partial circularity: the tailor is selected to score well on those metrics, so reporting them as proof of system effectiveness is not fully independent. However, the paper also evaluates on external downstream tasks (BBH, MMLU, Commonsense) that are not part of the objective, and the hardware accelerator, DVFS, and MoE router are assessed separately. No load-bearing result is justified solely by a self-citation, and no uniqueness theorem is imported from the authors' prior work. The headline 11.92x/7.36x numbers rest on a simulated accelerator and a slow FlexGen baseline, but those are correctness/reproducibility risks, not circularity. Overall the derivation chain is not forced by construction; score 2 is appropriate for the partial objective/evaluation overlap.
Assumptions & free parameters
free parameters (8)
- alpha and beta penalty exponents =
set to 2
- Latency and energy budgets T and E =
not specified in paper
- top-K starting points =
25
- Gradient step size eta =
0.8
- LoRA rank r and scaling alpha =
8 and 16
- Encoder-decoder hidden sizes =
64/64/200
- DVFS controller MLP size =
two-layer, <1K parameters
- Power model lookup table entries =
measured power values
assumptions (4)
- domain assumption Post-layout simulation results are representative of fabricated chip behavior
- domain assumption The pruning surrogate trained on a finite set of ratio-score pairs generalizes to unseen layer configurations
- domain assumption Sentence-embedding cosine similarity is a valid proxy for which LoRA adapter improves a given request
- domain assumption The RL-based DVFS controller converges to a good policy without destabilizing latency
invented entities (1)
-
28nm CLONE accelerator with LPU and SFU
Cite this review
Pith. "Pith review of CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge." pith.science (2026). https://pith.science/paper/P74LQJEO
@misc{pith2026250602847,
author = {Pith},
title = {Pith review of: CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge},
year = {2026},
howpublished = {\url{https://pith.science/paper/P74LQJEO}},
note = {Machine review of arXiv:2506.02847}
}
read the original abstract
Deploying large language models (LLMs) on edge devices is crucial for delivering fast responses and ensuring data privacy. However, the limited storage, weight, and power of edge devices make it difficult to deploy LLM-powered applications. These devices must balance latency requirements with energy consumption and model accuracy. In this paper, we first quantify the challenges of deploying LLMs on off-the-shelf edge devices and then we present CLONE, an in-depth algorithm-hardware co-design at both the model- and system-level that intelligently integrates real-time, energy optimization while maintaining robust generality. In order to maximize the synergistic benefits of these algorithms in always-on and intermediate edge computing settings, we specialize in a 28nm scalable hardware accelerator system. We implement and extensively evaluate CLONE on two off-the-shelf edge platforms. Experiments show that CLONE effectively accelerates the inference process up to 11.92x, and saves energy up to 7.36x, while maintaining high-generation.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Rewind: A better way to search your digital life, 2024
Rewind AI. Rewind: A better way to search your digital life, 2024
2024
-
[2]
An open-source framework for autonomous soc design with analog block generation
Tutu Ajayi, Sumanth Kamineni, Yaswanth K Cherivi- rala, Morteza Fayazi, Kyumin Kwon, Mehdi Saligane, Shourya Gupta, Chien-Hen Chen, Dennis Sylvester, David Blaauw, et al. An open-source framework for autonomous soc design with analog block generation. In2020 IFIP/IEEE 28th International Conference on Very Large Scale Integration (VLSI-SOC), pages 141–
-
[3]
The illustrated gpt-2 (visualizing trans- former language models).Jalammar
Jay Alammar. The illustrated gpt-2 (visualizing trans- former language models).Jalammar. github. io. https://jalammar. github. io/illustrated-gpt2, 2019
2019
-
[4]
Del Mundo, Mohammad Rastegari, and Mehrdad Farajtabar
Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, Karen Khatamifard, Minsik Cho, Carlo C. Del Mundo, Mohammad Rastegari, and Mehrdad Farajtabar. LLM in a flash: Efficient large language model inference with limited memory.CoRR, abs/2312.11514, 2023
arXiv 2023
-
[5]
Claude: A family of ai models, 2024
Anthropic. Claude: A family of ai models, 2024
2024
-
[6]
Apple intelligence
Apple. Apple intelligence. https://www.apple. com/apple-intelligence/, 2024
2024
-
[7]
Apple siri: Virtual assistant, 2024
Apple. Apple siri: Virtual assistant, 2024
2024
-
[8]
Croci, Marcelo Gen- nari Do Nascimento, Torsten Hoefler, and James Hens- man
Saleh Ashkboos, Maximilian L. Croci, Marcelo Gen- nari Do Nascimento, Torsten Hoefler, and James Hens- man. Slicegpt: Compress large language models by deleting rows and columns.CoRR, abs/2401.15024, 2024
arXiv 2024
Show all 162 references
-
[9]
25.1 a fully synthesizable distributed and scalable all-digital ldo in 10nm cmos
Suyoung Bang, Wootaek Lim, Charles Augustine, An- dres Malavasi, Muhammad Khellah, James Tschanz, and Vivek De. 25.1 a fully synthesizable distributed and scalable all-digital ldo in 10nm cmos. In2020 IEEE International Solid-State Circuits Conference- (ISSCC), pages 380–382. ...
2020
-
[10]
{NeuOS}: A {Latency- Predictable}{Multi-Dimensional} optimization frame- work for {DNN-driven} autonomous systems
Soroush Bateni and Cong Liu. {NeuOS}: A {Latency- Predictable}{Multi-Dimensional} optimization frame- work for {DNN-driven} autonomous systems. In2020 USENIX Annual Technical Conference (USENIX ATC 20), pages 371–385, 2020
2020
-
[11]
Bender, Timnit Gebru, Angelina McMillan- Major, and Shmargaret Shmitchell
Emily M. Bender, Timnit Gebru, Angelina McMillan- Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Madeleine Clare Elish, William Isaac, and Richard S. Zemel, editors,FAccT ’21: 2021 ACM Con- ference on Fairness, Accoun...
2021
-
[12]
PIQA: reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jian- feng Gao, and Yejin Choi. PIQA: reasoning about physical commonsense in natural language. InThe Thirty-Fourth AAAI Conference on Artificial Intelli- gence, AAAI 2020, The Thirty-Second Innovative Appli- cations of Artificial In...
2020
-
[13]
Hudson, Ehsan Adeli, Russ B
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ B. Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosse- lut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri S. Chat- terji, Annie S. Chen, Kathlee...
2021 arXiv
-
[14]
An estimate of an upper bound for the entropy of english
Peter F Brown, Stephen A Della Pietra, Vincent J Della Pietra, Jennifer C Lai, and Robert L Mercer. An estimate of an upper bound for the entropy of english. Computational Linguistics, 18(1):31–40, 1992
1992
-
[15]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Nee- lakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Je...
2020
-
[16]
A fre- quency compensation scheme for ldo voltage regula- tors.IEEE Transactions on Circuits and Systems I: Regular Papers, 51(6):1041–1050, 2004
Chaitanya K Chava and José Silva-Martínez. A fre- quency compensation scheme for ldo voltage regula- tors.IEEE Transactions on Circuits and Systems I: Regular Papers, 51(6):1041–1050, 2004
2004
-
[17]
Bge m3-embedding: Multi- lingual, multi-functionality, multi-granularity text em- beddings through self-knowledge distillation, 2024
Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. Bge m3-embedding: Multi- lingual, multi-functionality, multi-granularity text em- beddings through self-knowledge distillation, 2024
2024
-
[18]
Op- timizing energy efficiency of browsers in energy-aware scheduling-enabled mobile devices
Yonghun Choi, Seonghoon Park, and Hojung Cha. Op- timizing energy efficiency of browsers in energy-aware scheduling-enabled mobile devices. In Stephen A. Brewster, Geraldine Fitzpatrick, Anna L. Cox, and Vas- silis Kostakos, editors,The 25th Annual International Conference on ...
2019
-
[19]
Aakanksha Chowdhery, Sharan Narang, Jacob De- vlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Se- bastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, ...
2023
-
[20]
ML.ENERGY leaderboard
Jae-Won Chung, Jiachen Liu, Zhiyu Wu, Yuxuan Xia, and Mosharaf Chowdhury. ML.ENERGY leaderboard. https://ml.energy/leaderboard, 2023
2023
-
[21]
Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche, Eliza Ruther- ford, Tom Hennigan, Matthew J
Aidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake A. Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche, Eliza Ruther- ford, Tom Hennigan, Matthew J. Johnson, Albin Cas- sirer, Chris Jones, Elena...
2022
-
[22]
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. In Jill Burstein, Christy Doran, and Thamar Solorio, editors,Proceedings of the 2019 Con- ference of...
2019
-
[23]
Think you have solved ques- tion answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved ques- tion answering? try arc, the ai2 reasoning challenge. arXiv:1803.05457v1, 2018
2018 arXiv
-
[24]
Track and reduce co2 emissions from your computing.https://codecarbon.io/, 2023
codecarbon. Track and reduce co2 emissions from your computing.https://codecarbon.io/, 2023
2023
-
[25]
Llama 2: Inferencing on a single gpu
Dell. Llama 2: Inferencing on a single gpu. https: //infohub.delltechnologies.com/zh-cn/t/ llama-2-inferencing-on-a-single-gpu/, 2023
2023
-
[26]
Llm.int8(): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. Llm.int8(): 8-bit matrix multiplication for transformers at scale. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors,Advances in Neural Information Processing Systems 35: Annua...
2022
-
[27]
Pruner- zero: Evolving symbolic pruning metric from scratch for large language models
Peijie Dong, Lujun Li, Zhenheng Tang, Xiang Liu, Xinglin Pan, Qiang Wang, and Xiaowen Chu. Pruner- zero: Evolving symbolic pruning metric from scratch for large language models. InForty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27,...
2024
-
[28]
User-aware frame rate man- agement in android smartphones.ACM Trans
Begum Egilmez, Matthew Schuchhardt, Gokhan Memik, Raid Ayoub, Niranjan Soundararajan, and Michael Kishinevsky. User-aware frame rate man- agement in android smartphones.ACM Trans. Embed. Comput. Syst., 16(5s):131:1–131:17, 2017
2017
-
[29]
Bradley Chen, and Margo I
Yasuhiro Endo, Zheng Wang, J. Bradley Chen, and Margo I. Seltzer. Using latency to evaluate inter- active system performance. In Karin Petersen and Willy Zwaenepoel, editors,Proceedings of the Second USENIX Symposium on Operating Systems Design and Implementation (OSDI), Seatt...
1996
-
[30]
Depgraph: Towards any struc- tural pruning
Gongfan Fang, Xinyin Ma, Mingli Song, Michael Bi Mi, and Xinchao Wang. Depgraph: Towards any struc- tural pruning. InIEEE/CVF Conference on Com- puter Vision and Pattern Recognition, CVPR 2023, Van- couver, BC, Canada, June 17-24, 2023, pages 16091– 16101. IEEE, 2023
2023
-
[31]
Gpt-3: Its nature, scope, limits, and consequences.Minds and Machines, 30:681–694, 2020
Luciano Floridi and Massimo Chiriatti. Gpt-3: Its nature, scope, limits, and consequences.Minds and Machines, 30:681–694, 2020
2020
-
[32]
Sparsegpt: Massive language models can be accurately pruned in one- shot
Elias Frantar and Dan Alistarh. Sparsegpt: Massive language models can be accurately pruned in one- shot. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors,International Conference on Machine Learning, ICML 2023,...
2023
-
[33]
Beam search strategies for neural machine translation
Markus Freitag and Yaser Al-Onaizan. Beam search strategies for neural machine translation. In Thang Lu- ong, Alexandra Birch, Graham Neubig, and Andrew M. Finch, editors,Proceedings of the First Workshop on Neural Machine Translation, NMT@ACL 2017, Van- couver, Canada, August...
2017
-
[34]
Drive like a human: Rethinking autonomous driving with large language models
Daocheng Fu, Xin Li, Licheng Wen, Min Dou, Pinlong Cai, Botian Shi, and Yu Qiao. Drive like a human: Rethinking autonomous driving with large language models. InProceedings of the IEEE/CVF Winter Con- ference on Applications of Computer Vision, pages 910– 919, 2024
2024
-
[35]
Openllama: An open reproduction of llama
Xinyang Geng and Hao Liu. Openllama: An open reproduction of llama. https://github.com/ openlm-research/open_llama, May 2023
2023
-
[36]
Mahoney, and Kurt Keutzer
Amir Gholami, Zhewei Yao, Sehoon Kim, Coleman Hooper, Michael W. Mahoney, and Kurt Keutzer. AI and memory wall.IEEE Micro, 44(3):33–39, 2024
2024
-
[37]
GitHub Copilot: Your AI pair programmer
github. GitHub Copilot: Your AI pair programmer. https://github.com/features/copilot
-
[38]
Goodfellow, Yoshua Bengio, and Aaron C
Ian J. Goodfellow, Yoshua Bengio, and Aaron C. Courville.Deep Learning. Adaptive computation and machine learning. MIT Press, 2016
2016
-
[39]
Google assistant, 2024
Google. Google assistant, 2024
2024
-
[40]
Ml kit smart reply, 2024
Google. Ml kit smart reply, 2024
2024
-
[41]
Minillm: Knowledge distillation of large language models
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. Minillm: Knowledge distillation of large language models. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024
2024
-
[42]
Mistify: Au- tomating DNN model porting for on-device inference at the edge
Peizhen Guo, Bo Hu, and Wenjun Hu. Mistify: Au- tomating DNN model porting for on-device inference at the edge. In James Mickens and Renata Teixeira, ed- itors,18th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2021, April 12-14, 2021, pages 705–719. US...
2021
-
[43]
A reconfigurable floating- point compute-in-memory with analog exponent pre- processes.IEEE Solid-State Circuits Letters, 2024
Pengyu He, Yuanzhe Zhao, Heng Xie, Yang Wang, Shouyi Yin, Li Li, Yan Zhu, Rui P Martins, Chi-Hang Chan, and Minglei Zhang. A reconfigurable floating- point compute-in-memory with analog exponent pre- processes.IEEE Solid-State Circuits Letters, 2024
2024
-
[44]
A 28nm 314.6 tl- fops/w reconfigurable floating-point analog compute- in-memory macro with exponent approximation and two-stage sharing td-adc
Pengyu He, Yuanzhe Zhao, Heng Xie, Yang Wang, Shouyi Yin, Li Li, Yan Zhu, Rui Paulo Martins, Chi- Hang Chan, and Minglei Zhang. A 28nm 314.6 tl- fops/w reconfigurable floating-point analog compute- in-memory macro with exponent approximation and two-stage sharing td-adc. In202...
2024
-
[45]
Channel pruning for accelerating very deep neural networks
Yihui He, Xiangyu Zhang, and Jian Sun. Channel pruning for accelerating very deep neural networks. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017, pages 1398–1406. IEEE Computer Society, 2017
2017
-
[46]
Measuring massive multitask language under- standing.Proceedings of the International Conference on Learning Representations (ICLR), 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Stein- hardt. Measuring massive multitask language under- standing.Proceedings of the International Conference on Learning Representations (ICLR), 2021
2021
-
[47]
Variation- aware dynamic voltage/frequency scaling
Sebastian Herbert and Diana Marculescu. Variation- aware dynamic voltage/frequency scaling. In15th International Conference on High-Performance Com- puter Architecture (HPCA-15 2009), 14-18 February 2009, Raleigh, North Carolina, USA, pages 301–312. IEEE Computer Society, 2009
2009
-
[48]
Enright Jerger
Robert Hesse and Natalie D. Enright Jerger. Improving DVFS in nocs with coherence prediction. In André Ivanov, Diana Marculescu, Partha Pratim Pande, José Flich, and Karthik Pattabiraman, editors,Proceedings of the 9th International Symposium on Networks-on- Chip, NOCS 2015, V...
2015
-
[49]
Long short- term memory.Neural Comput., 9(8):1735–1780, 1997
Sepp Hochreiter and Jürgen Schmidhuber. Long short- term memory.Neural Comput., 9(8):1735–1780, 1997
1997
-
[50]
The curious case of neural text degener- ation
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degener- ation. In8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020
2020
-
[51]
An all-digital phase-locked loop (adpll)-based clock re- covery circuit.IEEE Journal of Solid-State Circuits, 34(8):1063–1073, 1999
Terng-Yin Hsu, Bai-Jue Shieh, and Chen-Yi Lee. An all-digital phase-locked loop (adpll)-based clock re- covery circuit.IEEE Journal of Solid-State Circuits, 34(8):1063–1073, 1999
1999
-
[52]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-Rank Adaptation of Large Language Models.arXiv e-prints, page arXiv:2106.09685, June 2021
2021 arXiv
-
[53]
CAMA: energy and memory efficient automata pro- cessing in content-addressable memories
Yi Huang, Zhiyu Chen, Dai Li, and Kaiyuan Yang. CAMA: energy and memory efficient automata pro- cessing in content-addressable memories. InIEEE International Symposium on High-Performance Com- puter Architecture, HPCA 2022, Seoul, South Korea, April 2-6, 2022, pages 25–37. IEEE, 2022
2022
-
[54]
New solutions on llm acceleration, optimization, and application
Yingbing Huang, Lily Jiaxin Wan, Hanchen Ye, Manvi Jha, Jinghua Wang, Yuhong Li, Xiaofan Zhang, and Deming Chen. New solutions on llm acceleration, optimization, and application. InProceedings of the 61st ACM/IEEE Design Automation Conference, pages 1–4, 2024
2024
-
[55]
Quantized neural networks: Training neural networks with low precision weights and activations.Journal of Machine Learning Research, 18(187):1–30, 2018
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. Quantized neural networks: Training neural networks with low precision weights and activations.Journal of Machine Learning Research, 18(187):1–30, 2018
2018
-
[56]
Transformers documentation, 2024
HuggingFace Team. Transformers documentation, 2024
2024
-
[57]
Just-in-time quan- tization with processing-in-memory for efficient ml training, 2023
Mohamed Assem Ibrahim, Shaizeen Aga, Ada Li, Su- chita Pati, and Mahzabeen Islam. Just-in-time quan- tization with processing-in-memory for efficient ml training, 2023
2023
-
[58]
Mem: Your ai-powered assistant, 2024
Mem Inc. Mem: Your ai-powered assistant, 2024
2024
-
[59]
Jacobs, Michael I
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton. Adaptive mixtures of local experts.Neural Comput., 3(1):79–87, 1991
1991
-
[60]
Llm performance pre- dictors are good initializers for architecture search
Ganesh Jawahar, Muhammad Abdul-Mageed, Laks VS Lakshmanan, and Dujian Ding. Llm performance pre- dictors are good initializers for architecture search. arXiv preprint arXiv:2310.16712, 2023
2023 arXiv
-
[61]
Robot control using llama: Bridging ai and robotics
Kabilankb. Robot control using llama: Bridging ai and robotics. Medium, 2024. Oct 13, 2024
2024
-
[62]
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L Littman, and An- drew W Moore. Reinforcement learning: A survey. Journal of artificial intelligence research, 4:237–285, 1996
1996
-
[63]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models.CoRR, abs/2001.08361, 2020
2001 arXiv
-
[64]
The breakthrough memory solutions for improved performance on LLM inference.IEEE Micro, 44(3):40–48, 2024
Byeongho Kim, Sanghoon Cha, Sangsoo Park, Jieun Lee, Sukhan Lee, Shinhaeng Kang, Jinin So, Kyungsoo Kim, Jin Jung, Jong-Geon Lee, Sunjung Lee, Yoonah Paik, Hyeonsu Kim, Jin-Seong Kim, Won-Jo Lee, Yuh- wan Ro, Yeongon Cho, Jin Hyun Kim, Joon-Ho Song, Jaehoon Yu, Seungwon Lee, J...
2024
-
[65]
Wonyoung Kim, Meeta Sharma Gupta, Gu-Yeon Wei, and David M. Brooks. System level analysis of fast, per-core DVFS using on-chip switching regulators. In 14th International Conference on High-Performance Computer Architecture (HPCA-14 2008), 16-20 Febru- ary 2008, Salt Lake City...
2008
-
[66]
Enhancing energy efficiency of multimedia applications in heterogeneous mobile multi-core pro- cessors.IEEE Trans
Young Geun Kim, Minyong Kim, and Sung Woo Chung. Enhancing energy efficiency of multimedia applications in heterogeneous mobile multi-core pro- cessors.IEEE Trans. Computers, 66(11):1878–1889, 2017
2017
-
[67]
A novel gpu power model for accurate smartphone power break- down.ETRI journal, 37(1):157–164, 2015
Young Geun Kim, Minyong Kim, Jae Min Kim, Miny- oung Sung, and Sung Woo Chung. A novel gpu power model for accurate smartphone power break- down.ETRI journal, 37(1):157–164, 2015
2015
-
[68]
A survey on recent os-level energy management tech- niques for mobile processing units.IEEE Trans
Young Geun Kim, Joonho Kong, and Sung Woo Chung. A survey on recent os-level energy management tech- niques for mobile processing units.IEEE Trans. Paral- lel Distributed Syst., 29(10):2388–2401, 2018
2018
-
[69]
Autoscale: Energy efficiency optimization for stochastic edge in- ference using reinforcement learning
Young Geun Kim and Carole-Jean Wu. Autoscale: Energy efficiency optimization for stochastic edge in- ference using reinforcement learning. In53rd Annual IEEE/ACM International Symposium on Microarchi- tecture, MICRO 2020, Athens, Greece, October 17-21, 2020, pages 1082–1096. I...
2020
-
[70]
InProceedings of the Fourteenth EuroSys Conference 2019, pages 1–15, 2019
Youngsok Kim, Joonsung Kim, Dongju Chae, Daehyun Kim, and Jangwoo Kim.µlayer: Low latency on-device inference using cooperative single-layer acceleration and processor-friendly quantization. InProceedings of the Fourteenth EuroSys Conference 2019, pages 1–15, 2019
2019
-
[71]
Distillm: Towards streamlined distillation for large language models
Jongwoo Ko, Sungnyun Kim, Tianyi Chen, and Se- Young Yun. Distillm: Towards streamlined distillation for large language models. InForty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024
2024
-
[72]
Resnet 50.Convolu- tional neural networks with swift for tensorflow: image recognition and dataset categorization, pages 63–72, 2021
Brett Koonce and Brett Koonce. Resnet 50.Convolu- tional neural networks with swift for tensorflow: image recognition and dataset categorization, pages 63–72, 2021
2021
-
[73]
Automatic domain-specific soc design for autonomous unmanned aerial vehicles
Srivatsan Krishnan, Zishen Wan, Kshitij Bhardwaj, Paul Whatmough, Aleksandra Faust, Sabrina Neuman, Gu-Yeon Wei, David Brooks, and Vijay Janapa Reddi. Automatic domain-specific soc design for autonomous unmanned aerial vehicles. In2022 55th IEEE/ACM International Symposium on ...
2022
-
[75]
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gon- zalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Jason Flinn, Margo I. Seltzer, Pe- ter Druschel, Antoine Kaufmann,...
2023
-
[76]
Deep learning.nature, 521(7553):436–444, 2015
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning.nature, 521(7553):436–444, 2015
2015
-
[77]
On-chip memory technology design space explorations for mobile deep neural network ac- celerators
Haitong Li, Mudit Bhargava, Paul N Whatmough, and H-S Philip Wong. On-chip memory technology design space explorations for mobile deep neural network ac- celerators. InProceedings of the 56th Annual Design Automation Conference 2019, pages 1–6, 2019
2019
-
[78]
Energydx: Diag- nosing energy anomaly in mobile apps by identifying the manifestation point
Li Li, Xiaorui Wang, and Feng Qin. Energydx: Diag- nosing energy anomaly in mobile apps by identifying the manifestation point. In2020 IEEE 40th Interna- tional Conference on Distributed Computing Systems (ICDCS), pages 256–266. IEEE, 2020
2020
-
[79]
Deep reinforcement learning: An overview
Yuxi Li. Deep reinforcement learning: An overview. arXiv preprint arXiv:1701.07274, 2017
2017 arXiv
-
[80]
From system 1 to system 2: A survey of reasoning large language models, 2025
Zhong-Zhi Li, Duzhen Zhang, Ming-Liang Zhang, Ji- axin Zhang, Zengyan Liu, Yuxuan Yao, Haotian Xu, Junhao Zheng, Pei-Jie Wang, Xiuyi Chen, Yingying Zhang, Fei Yin, Jiahua Dong, Zhiwei Li, Bao-Long Bi, Ling-Rui Mei, Junfeng Fang, Zhijiang Guo, Le Song, and Cheng-Lin Liu. From s...
2025
-
[81]
BAT: behavior-aware human-like trajectory prediction for autonomous driving.CoRR, abs/2312.06371, 2023
Haicheng Liao, Zhenning Li, Huanming Shen, Wenx- uan Zeng, Dongping Liao, Guofa Li, Shengbo Eben Li, and Chengzhong Xu. BAT: behavior-aware human-like trajectory prediction for autonomous driving.CoRR, abs/2312.06371, 2023
2023 arXiv
-
[82]
Edgar Liberis and Nicholas D Lane. Differentiable neural network pruning to enable smart applications on microcontrollers.Proceedings of the ACM on Inter- active, Mobile, Wearable and Ubiquitous Technologies, 6(4):1–19, 2023
2023
-
[83]
Par- rot: Efficient serving of llm-based applications with semantic variable
Chaofan Lin, Zhenhua Han, Chengruidong Zhang, Yuqing Yang, Fan Yang, Chen Chen, and Lili Qiu. Par- rot: Efficient serving of llm-based applications with semantic variable. In Ada Gavrilovska and Douglas B. Terry, editors,18th USENIX Symposium on Operat- ing Systems Design and ...
2024
-
[84]
AWQ: activation-aware weight quantization for LLM compression and acceler- ation.CoRR, abs/2306.00978, 2023
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. AWQ: activation-aware weight quantization for LLM compression and acceler- ation.CoRR, abs/2306.00978, 2023
2023 arXiv
-
[85]
A rein- forcement learning-based power management frame- work for green computing data centers
Xue Lin, Yanzhi Wang, and Massoud Pedram. A rein- forcement learning-based power management frame- work for green computing data centers. In2016 IEEE International Conference on Cloud Engineering, IC2E 2016, Berlin, Germany, April 4-8, 2016, pages 135–138. IEEE Computer Society, 2016
2016
-
[86]
A weak puf-assisted strong PUF with inherent immunity to modeling attacks and ultra-low BER.IEEE Trans
Jiahao Liu, Yuanzhe Zhao, Yan Zhu, Chi-Hang Chan, and Rui Paulo Martins. A weak puf-assisted strong PUF with inherent immunity to modeling attacks and ultra-low BER.IEEE Trans. Circuits Syst. I Regul. Pap., 69(12):4898–4907, 2022
2022
-
[87]
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing.ACM Comput
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing.ACM Comput. Surv., 55(9):195:1–195:35, 2023
2023
-
[88]
Optimizing llm queries in relational work- loads.arXiv preprint arXiv:2403.05821, 2024
Shu Liu, Asim Biswal, Audrey Cheng, Xiangxi Mo, Shiyi Cao, Joseph E Gonzalez, Ion Stoica, and Matei Zaharia. Optimizing llm queries in relational work- loads.arXiv preprint arXiv:2403.05821, 2024
2024 arXiv
-
[89]
P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks.CoRR, abs/2110.07602, 2021
Xiao Liu, Kaixuan Ji, Yicheng Fu, Zhengxiao Du, Zhilin Yang, and Jie Tang. P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks.CoRR, abs/2110.07602, 2021
2021 arXiv
-
[90]
S2ta: Exploiting structured sparsity for energy-efficient mobile cnn acceleration
Zhi-Gang Liu, Paul N Whatmough, Yuhao Zhu, and Matthew Mattina. S2ta: Exploiting structured sparsity for energy-efficient mobile cnn acceleration. In2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA), pages 573–586. IEEE, 2022
2022
-
[91]
Deja vu: Contextual sparsity for efficient llms at inference time
Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Ré, and Beidi Chen. Deja vu: Contextual sparsity for efficient llms at inference time. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara En...
2023
-
[92]
Prediction-guided performance-energy trade-off for in- teractive applications
Daniel Lo, Taejoon Song, and G Edward Suh. Prediction-guided performance-energy trade-off for in- teractive applications. InProceedings of the 48th In- ternational Symposium on Microarchitecture, pages 508–520, 2015
2015
-
[93]
Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. Building a large annotated corpus of english: The penn treebank.Comput. Linguistics, 19(2):313–330, 1993
1993
-
[94]
Shortgpt: Layers in large language mod- els are more redundant than you expect.CoRR, abs/2403.03853, 2024
Xin Men, Mingyu Xu, Qingyu Zhang, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, and Weipeng Chen. Shortgpt: Layers in large language mod- els are more redundant than you expect.CoRR, abs/2403.03853, 2024
2024 arXiv
-
[95]
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. In 5th International Conference on Learning Representa- tions, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017
2017
-
[96]
Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Aakanksha Chowdh- ery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Amé...
2024 arXiv
-
[97]
Your Everyday AI Companion | Microsoft Bing.https://www.bing.com/new
microsoft. Your Everyday AI Companion | Microsoft Bing.https://www.bing.com/new
-
[98]
Azure cognitive services - text analytics: Smart reply, 2024
Microsoft. Azure cognitive services - text analytics: Smart reply, 2024
2024
-
[99]
Deepspeed: Advancing the science of ai through efficient training of large models, 2024
Microsoft DeepSpeed Team. Deepspeed: Advancing the science of ai through efficient training of large models, 2024
2024
-
[100]
Can a suit of armor conduct electricity? A new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? A new dataset for open book question answering. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii, editors,Proceedings of the 2018 Con- ference on Empir...
2018
-
[101]
Human-level control through deep reinforcement learning.nature, 518(7540):529– 533, 2015
V olodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. Human-level control through deep reinforcement learning.nature, 518(7540):529– 533, 2015
2015
-
[102]
Importance estimation for neu- ral network pruning
Pavlo Molchanov, Arun Mallya, Stephen Tyree, Iuri Frosio, and Jan Kautz. Importance estimation for neu- ral network pruning. InIEEE Conference on Com- puter Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 11264– 11272. Computer Vision Fo...
2019
-
[103]
Pruning convolutional neural net- works for resource efficient inference.arXiv preprint arXiv:1611.06440, 2016
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz. Pruning convolutional neural net- works for resource efficient inference.arXiv preprint arXiv:1611.06440, 2016
2016 arXiv
-
[104]
Nvidia jetson nano, 2023
nano. Nvidia jetson nano, 2023
2023
-
[105]
Carpenter, and Magnus Själander
Rajiv Nishtala, Vinicius Petrucci, Paul M. Carpenter, and Magnus Själander. Twig: Multi-agent task manage- ment for colocated latency-critical cloud services. In IEEE International Symposium on High Performance Computer Architecture, HPCA 2020, San Diego, CA, USA, February 22-...
2020
-
[106]
Nvidia jetson orin - autonomous ma- chines - nvidia
NVIDIA. Nvidia jetson orin - autonomous ma- chines - nvidia. https://www.nvidia.com/en-us/ autonomous-machines/embedded-systems/ jetson-orin/, 2024
2024
-
[107]
Exegpt: Constraint-aware resource scheduling for LLM inference
Hyungjun Oh, Kihong Kim, Jaemin Kim, Sungkyun Kim, Junyeol Lee, Du-seong Chang, and Jiwon Seo. Exegpt: Constraint-aware resource scheduling for LLM inference. In Rajiv Gupta, Nael B. Abu- Ghazaleh, Madan Musuvathi, and Dan Tsafrir, editors, Proceedings of the 29th ACM Internat...
2024
-
[108]
Chatgpt, 2022
OpenAI. Chatgpt, 2022
2022
-
[109]
GPT-4 technical report.CoRR, abs/2303.08774, 2023
OpenAI. GPT-4 technical report.CoRR, abs/2303.08774, 2023
2023 arXiv
-
[110]
Otter: Transcription and note-taking with ai, 2024
Otter.ai. Otter: Transcription and note-taking with ai, 2024
2024
-
[111]
Splitwise: Efficient generative LLM inference using phase splitting.CoRR, abs/2311.18677, 2023
Pratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri, Aashaka Shah, Saeed Maleki, and Ricardo Bian- chini. Splitwise: Efficient generative LLM inference using phase splitting.CoRR, abs/2311.18677, 2023
2023 arXiv
-
[112]
Expanding datacenter capacity with DVFS boosting: A safe and scalable deployment experience
Leonardo Piga, Iyswarya Narayanan, Aditya Sundar- rajan, Matt Skach, Qingyuan Deng, Biswadip Maity, Manoj Chakkaravarthy, Alison Huang, Abhishek Dhan- otia, and Parth Malani. Expanding datacenter capacity with DVFS boosting: A safe and scalable deployment experience. In Rajiv ...
2024
-
[113]
Pytorch: An open source ma- chine learning framework, 2024
PyTorch Contributors. Pytorch: An open source ma- chine learning framework, 2024
2024
-
[114]
Llama-v2-7B-Chat Quantized Model
Qualcomm Technologies, Inc. Llama-v2-7B-Chat Quantized Model. https://aihub.qualcomm. com/models/llama_v2_7b_chat_quantized, 2024. Qualcomm AI Hub, 2024
2024
-
[115]
Language mod- els are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language mod- els are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019
2019
-
[116]
Winogrande: An adversarial winograd schema challenge at scale.arXiv preprint arXiv:1907.10641, 2019
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhaga- vatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale.arXiv preprint arXiv:1907.10641, 2019
1907 arXiv
-
[117]
Le, Geoffrey E
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V . Le, Geoffrey E. Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In5th Interna- tional Conference on Learning Representations, ICLR 2017, Toulon, Fra...
2017
-
[118]
Flexgen: High- throughput generative inference of large language mod- els with a single gpu
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christo- pher Ré, Ion Stoica, and Ce Zhang. Flexgen: High- throughput generative inference of large language mod- els with a single gpu. InInternational Conference on Machine Learning, ...
2023
-
[119]
Progprompt: Gen- erating situated robot task plans using large language models
Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg. Progprompt: Gen- erating situated robot task plans using large language models. InIEEE International Conference on Robotics and Automation, I...
2023
-
[120]
Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mah- davi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Kumar Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Nathaneal Schärli, Aakanksha Chowdh- ery, Philip Andrew Mansf...
2022 arXiv
-
[121]
Llm- planner: Few-shot grounded planning for embodied agents with large language models
Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M Sadler, Wei-Lun Chao, and Yu Su. Llm- planner: Few-shot grounded planning for embodied agents with large language models. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 2998–3009, 2023
2023
-
[122]
Powerinfer: Fast large language model serving with a consumer-grade gpu
Yixin Song, Zeyu Mi, Haotong Xie, and Haibo Chen. Powerinfer: Fast large language model serving with a consumer-grade gpu. InProceedings of the ACM SIGOPS 30th Symposium on Operating Systems Prin- ciples, pages 590–606, 2024
2024
-
[123]
Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, Ali...
2022 arXiv
-
[124]
Ai index report
stanford. Ai index report. https://aiindex. stanford.edu/report/, 2024. Accessed: 2024-07
2024
-
[125]
Mobile ram usage worldwide from 1q-19 to 1q-21 (in gb per device)
Statista Inc. Mobile ram usage worldwide from 1q-19 to 1q-21 (in gb per device). www.statista.com/statistics/1057679/mobile-ram- usage-worldwide-by-average-size-per-device/, 2021
2021
-
[126]
Fedhybrid: Breaking the memory wall of federated learning via hybrid tensor manage- ment
Kahou Tam, Chunlin Tian, Li Li, Haikai Zhao, and ChengZhong Xu. Fedhybrid: Breaking the memory wall of federated learning via hybrid tensor manage- ment. InProceedings of the 22nd ACM Conference on Embedded Networked Sensor Systems, pages 394–408, 2024
2024
-
[127]
Edgebert: Sentence-level energy optimizations for latency-aware multi-task nlp inference
Thierry Tambe, Coleman Hooper, Lillian Pentecost, Tianyu Jia, En-Yu Yang, Marco Donato, Victor Sanh, Paul Whatmough, Alexander M Rush, David Brooks, et al. Edgebert: Sentence-level energy optimizations for latency-aware multi-task nlp inference. InMICRO- 54: 54th Annual IEEE/A...
2021
-
[128]
Tensorflow: An open source machine learning framework for everyone, 2024
TensorFlow Contributors. Tensorflow: An open source machine learning framework for everyone, 2024
2024
-
[129]
Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023
Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting. Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023
1930
-
[130]
Harmony: Heterogeneity-aware hi- erarchical management for federated learning system
Chunlin Tian, Li Li, Zhan Shi, Jun Wang, and ChengZhong Xu. Harmony: Heterogeneity-aware hi- erarchical management for federated learning system. In2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO), pages 631–645. IEEE, 2022
2022
-
[131]
Hydralora: An asymmetric lora architecture for efficient fine-tuning.Advances in Neu- ral Information Processing Systems, 37:9565–9584, 2024
Chunlin Tian, Zhan Shi, Zhijiang Guo, Li Li, and Cheng-Zhong Xu. Hydralora: An asymmetric lora architecture for efficient fine-tuning.Advances in Neu- ral Information Processing Systems, 37:9565–9584, 2024
2024
-
[132]
Llama: Open and efficient foundation language models.CoRR, abs/2302.13971, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language...
2023 arXiv
-
[133]
Llama 2: Open foundation and fine-tuned chat models.CoRR, abs/2307.09288, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Al- bert, Amjad Almahairi, Yasmine Babaei, Nikolay Bash- lykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy...
2023 arXiv
-
[134]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V . N. Vishwanathan, and Roman Garnett,...
2017
-
[135]
Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M
Jason Wei, Maarten Bosma, Vincent Y . Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V . Le. Finetuned language models are zero-shot learners.CoRR, abs/2109.01652, 2021
2021 arXiv
-
[136]
Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent abilities of large language models.Tr...
2022
-
[137]
Outlier suppression+: Accurate quantization of large language models by equivalent and effective shifting and scaling
Xiuying Wei, Yunchen Zhang, Yuhang Li, Xiangguo Zhang, Ruihao Gong, Jinyang Guo, and Xianglong Liu. Outlier suppression+: Accurate quantization of large language models by equivalent and effective shifting and scaling. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proc...
2023
-
[138]
Autodroid: Llm- powered task automation in android
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. Autodroid: Llm- powered task automation in android. InProceedings of the 30th Annual International Conference on Mobile Computing and Networking (Mob...
2024
-
[139]
Machine learn- ing at facebook: Understanding inference at the edge
Carole-Jean Wu, David Brooks, Kevin Chen, Douglas Chen, Sy Choudhury, Marat Dukhan, Kim Hazelwood, Eldad Isaac, Yangqing Jia, Bill Jia, et al. Machine learn- ing at facebook: Understanding inference at the edge. In2019 IEEE international symposium on high perfor- mance compute...
2019
-
[140]
Heterogeneity- aware memory efficient federated learning via progres- sive layer freezing
Yebo Wu, Li Li, Chunlin Tian, Tao Chang, Chi Lin, Cong Wang, and Cheng-Zhong Xu. Heterogeneity- aware memory efficient federated learning via progres- sive layer freezing. In2024 IEEE/ACM 32nd Inter- national Symposium on Quality of Service (IWQoS), pages 1–10. IEEE, 2024
-
[142]
A survey on fed- erated fine-tuning of large language models.CoRR, abs/2503.12016, 2025
Yebo Wu, Chunlin Tian, Jingguang Li, He Sun, Kahou Tam, Li Li, and Chengzhong Xu. A survey on fed- erated fine-tuning of large language models.CoRR, abs/2503.12016, 2025
2025
-
[143]
Smoothquant: Accu- rate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu, Julien Demouth, and Song Han. Smoothquant: Accu- rate and efficient post-training quantization for large language models. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, ...
2023
-
[144]
Llm4drive: A survey of large language models for autonomous driving.arXiv preprint arXiv:2311.01043, 2023
Zhenjie Yang, Xiaosong Jia, Hongyang Li, and Junchi Yan. Llm4drive: A survey of large language models for autonomous driving.arXiv preprint arXiv:2311.01043, 2023
2023 arXiv
-
[145]
Creating large language models on your laptop
Xinyu Ye, Zhe Wang, Haihao Shen, Yu Luo, and Han- wen Chang. Creating large language models on your laptop. Medium, 2023
2023
-
[146]
Phonelm: An efficient and capable small language model family through princi- pled pre-training.arXiv preprint arXiv:2411.05046, 2024
Rongjie Yi, Xiang Li, Weikai Xie, Zhenyan Lu, Chenghua Wang, Ao Zhou, Shangguang Wang, Xiwen Zhang, and Mengwei Xu. Phonelm: An efficient and capable small language model family through princi- pled pre-training.arXiv preprint arXiv:2411.05046, 2024
2024 arXiv
-
[147]
Elms: Elasticized large language models on mobile devices.arXiv preprint arXiv:2409.09071, 2024
Wangsong Yin, Rongjie Yi, Daliang Xu, Gang Huang, Mengwei Xu, and Xuanzhe Liu. Elms: Elasticized large language models on mobile devices.arXiv preprint arXiv:2409.09071, 2024
2024
-
[148]
Orca: A distributed serving system for transformer-based generative mod- els
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soo- jeong Kim, and Byung-Gon Chun. Orca: A distributed serving system for transformer-based generative mod- els. In Marcos K. Aguilera and Hakim Weather- spoon, editors,16th USENIX Symposium on Operat- ing Systems Design and Implem...
2022
-
[149]
Mobile foundation model as firmware.arXiv preprint arXiv:2308.14363, 2023
Jinliang Yuan, Chen Yang, Dongqi Cai, Shihe Wang, Xin Yuan, Zeling Zhang, Xiang Li, Dingge Zhang, Hanzi Mei, Xianqing Jia, et al. Mobile foundation model as firmware.arXiv preprint arXiv:2308.14363, 2023
2023 arXiv
-
[150]
Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning.CoRR, abs/2309.05444, 2023
Ted Zadouri, Ahmet Üstün, Arash Ahmadian, Beyza Er- mis, Acyr Locatelli, and Sara Hooker. Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning.CoRR, abs/2309.05444, 2023
2023 arXiv
-
[151]
Hellaswag: Can a machine re- ally finish your sentence? In Anna Korhonen, David R
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine re- ally finish your sentence? In Anna Korhonen, David R. Traum, and Lluís Màrquez, editors,Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL...
2019
-
[152]
Heterogeneity-aware coordination for feder- ated learning via stitching pre-trained blocks
Shichen Zhan, Yebo Wu, Chunlin Tian, Yan Zhao, and Li Li. Heterogeneity-aware coordination for feder- ated learning via stitching pre-trained blocks. In2024 IEEE/ACM 32nd International Symposium on Quality of Service (IWQoS), pages 1–10. IEEE, 2024
-
[153]
Felix: Optimizing tensor programs with gradient descent
Yifan Zhao, Hashim Sharif, Vikram Adve, and Sasa Misailovic. Felix: Optimizing tensor programs with gradient descent. InProceedings of the 29th ACM In- ternational Conference on Architectural Support for Programming Languages and Operating Systems, Vol- ume 3, pages 367–381, 2024
2024
-
[154]
Approxcaliper: A programmable framework for application-aware neural network optimization
Yifan Zhao, Hashim Sharif, Peter Pao-Huang, Vatsin Shah, Arun Narenthiran Sivakumar, Ma- teus Valverde Gasparino, Abdulrahman Mahmoud, Nathan Zhao, Sarita Adve, Girish Chowdhary, et al. Approxcaliper: A programmable framework for application-aware neural network optimization. ...
2023
-
[155]
A 28-nm 3.32- nj/frame compute-in-memory cnn processor with layer fusion for always-on applications.IEEE Transactions on Circuits and Systems I: Regular Papers, 2025
Yuanzhe Zhao, Pengyu He, Yan Zhu, Rui P Martins, Chi-Hang Chan, and Minglei Zhang. A 28-nm 3.32- nj/frame compute-in-memory cnn processor with layer fusion for always-on applications.IEEE Transactions on Circuits and Systems I: Regular Papers, 2025
2025
-
[156]
A 28nm value-wise hybrid-domain compute-in-memory macro with heterogeneous mem- ory fabric and asynchronous sparsity manager
Yuanzhe Zhao, Yang Wang, Yuheng Wang, Heng Xie, Yan Zhu, RP Martins, Chi-Hang Chan, Shouyi Yin, and Minglei Zhang. A 28nm value-wise hybrid-domain compute-in-memory macro with heterogeneous mem- ory fabric and asynchronous sparsity manager. In2025 IEEE Custom Integrated Circui...
2025
-
[157]
A reconfigurable 0.69-1.02 nj/classification biomedical ai processor for intelligent health monitoring devices
Yuanzhe Zhao, Yuheng Wang, Zijian Wang, Yan Zhu, RP Martins, Chi-Hang Chan, and Minglei Zhang. A reconfigurable 0.69-1.02 nj/classification biomedical ai processor for intelligent health monitoring devices. In2025 IEEE Custom Integrated Circuits Conference (CICC), pages 1–3. I...
2025
-
[158]
A one-shot floating-point compute-in- memory macro featuring pvt robustness and mismatch tolerance for edge llms
Yuanzhe Zhao, Heng Xie, Zijian Wang, Chunlin Tian, Li Li, Yan Zhu, RP Martins, Chi-Hang Chan, and Min- glei Zhang. A one-shot floating-point compute-in- memory macro featuring pvt robustness and mismatch tolerance for edge llms. In2025 IEEE Custom Inte- grated Circuits Confere...
2025
-
[159]
A double- mode sparse compute-in-memory macro with recon- figurable single and dual layer computation
Yuanzhe Zhao, Minglei Zhang, Pengyu He, Yan Zhu, Chi-Hang Chan, and Rui Paulo Martins. A double- mode sparse compute-in-memory macro with recon- figurable single and dual layer computation. InIEEE Custom Integrated Circuits Conference, CICC 2023, San Antonio, TX, USA, April 23...
2023
-
[160]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuo- han Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena. In Alice Oh, Tristan Nau- mann, Amir Gl...
2023
-
[161]
Event-based scheduling for energy-efficient qos (eqos) in mobile web applications
Yuhao Zhu, Matthew Halpern, and Vijay Janapa Reddi. Event-based scheduling for energy-efficient qos (eqos) in mobile web applications. In21st IEEE Interna- tional Symposium on High Performance Computer Ar- chitecture, HPCA 2015, Burlingame, CA, USA, Febru- ary 7-11, 2015, page...
2015
-
[162]
Ai robotics case - controlling robots with llms (large language models)
Ömer Bilgin Bilgili. Ai robotics case - controlling robots with llms (large language models). acrome,
-
[945]
USENIX Association, 2024
2024
-
[2024]
Acrome Robotics Blog, October 23, 2024
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.