Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:20:30.299592Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 100 of 162 outbound references and 1 inbound Pith citation observation for arXiv:2506.02847.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:20:30.299592Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-01T07:45:22.445395Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T07:45:29.205817Z
100 of 162 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3cc23a94-ad17-4c6c-92a8-25b94e6a0924 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Rewind: A better way to search your digital life, 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 897b39d5-ddef-4220-98f7-a17bea375f7f · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge An open-source framework for autonomous soc design with analog block generation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3fd116f-ad1c-453a-82d5-acdf9a3eced5 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge The illustrated gpt-2 (visualizing trans- former language models).Jalammar
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a21b22ec-6027-4945-98ae-68596982659f · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 292e48e6-24ab-42f0-8ff2-9c9d8af50a64 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Claude: A family of ai models, 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d336f0-d870-4a8c-bd09-1d635053d3fa · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Apple intelligence
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff3fe49a-584b-442e-8c18-875471b3ad65 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Apple siri: Virtual assistant, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c47924d-2585-4a07-af14-b99b5b35ee21 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge SliceGPT: Compress Large Language Models by Deleting Rows and Columns
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37fd1495-cc8f-4f20-9b83-914429bddfe3 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge 25.1 a fully synthesizable distributed and scalable all-digital ldo in 10nm cmos
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7146166-892c-4f84-bf31-4346bce7de12 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge {NeuOS}: A {Latency- Predictable}{Multi-Dimensional} optimization frame- work for {DNN-driven} autonomous systems
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18e0b18a-e656-4515-b77f-133a1022c919 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Bender, Timnit Gebru, Angelina McMillan- Major, and Shmargaret Shmitchell
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 095e4547-b476-4382-8ec6-99492013f2b9 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge PIQA: reasoning about physical commonsense in natural language
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ceb2af1-898b-4870-97cc-203e0d496c8a · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge On the Opportunities and Risks of Foundation Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da7612c2-af60-4652-a382-cd1498fb8143 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge An estimate of an upper bound for the entropy of english
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f5a3be4-0875-4749-9f67-3d1821e7fdb4 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e89f9a1-9fd3-4cf6-a3df-6948022d7d73 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge A fre- quency compensation scheme for ldo voltage regula- tors.IEEE Transactions on Circuits and Systems I: Regular Papers, 51(6):1041–1050, 2004
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3838fe1c-f4e0-4005-9923-c6954dca22e9 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Bge m3-embedding: Multi- lingual, multi-functionality, multi-granularity text em- beddings through self-knowledge distillation, 2024
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3ea6e24-12bb-45ea-8dfd-9c691b9bacf5 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Op- timizing energy efficiency of browsers in energy-aware scheduling-enabled mobile devices
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 872e40fe-bb1c-4730-8535-aacdb0088def · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77a962fe-6853-4879-b4d7-000581fb08d1 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge ML.ENERGY leaderboard
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45893bdb-be31-49d0-9e61-c067e8026f8c · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche, Eliza Ruther- ford, Tom Hennigan, Matthew J
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fc02c78-a21c-4b6f-aad3-4c110cdd1624 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Boolq: Exploring the surprising difficulty of natural yes/no questions
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64ffa5f7-1128-4cae-b93d-847c1d46ee0a · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b4ae27-2128-4ce2-a496-5b21c0ba3f82 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Track and reduce co2 emissions from your computing.https://codecarbon.io/, 2023
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 796f6d98-7a46-477d-a3bb-58d44d4844a3 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Llama 2: Inferencing on a single gpu
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 604a50a1-60b3-4f03-9137-e059d4a4020c · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Llm.int8(): 8-bit matrix multiplication for transformers at scale
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64368e54-f293-4069-ae4d-d84d825cc5ef · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Pruner- zero: Evolving symbolic pruning metric from scratch for large language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f2949f7-6d06-4034-a4f3-653ce46f9761 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge User-aware frame rate man- agement in android smartphones.ACM Trans
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 041a2a81-ab56-44b6-b9db-db536e6254ae · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Bradley Chen, and Margo I
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9583e905-eaa1-44e8-9835-575700d5f855 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Depgraph: Towards any struc- tural pruning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de9f1fce-d890-4ca5-b48d-1f823c8af701 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Gpt-3: Its nature, scope, limits, and consequences.Minds and Machines, 30:681–694, 2020
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb189c35-5f9d-45bd-9114-056fddd214de · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Sparsegpt: Massive language models can be accurately pruned in one- shot
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b55ce9d8-21c3-4720-9cd7-b87bd6aeaafd · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Beam search strategies for neural machine translation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69364cba-1976-45da-a39a-b3df3fead741 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Drive like a human: Rethinking autonomous driving with large language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c860af55-87db-4946-83fd-00f015302dff · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Openllama: An open reproduction of llama
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4d77d22-77ec-4e93-8ea0-00aa98e048e2 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Mahoney, and Kurt Keutzer
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f17d1f1f-d07a-4a6d-8309-55bc9119823b · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge GitHub Copilot: Your AI pair programmer
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f64bbe97-1e6f-42d7-84fa-8450a8f60850 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Goodfellow, Yoshua Bengio, and Aaron C
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b3382f0-4364-4c1c-9d32-74e10676352e · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Google assistant, 2024
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f2bac2c-c034-4d83-a651-89bceb96265b · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Ml kit smart reply, 2024
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cf2c75c-bcf0-4cef-ae89-6262d90269b6 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Minillm: Knowledge distillation of large language models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48647cfb-269f-4b74-8277-6b6b16bd0223 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Mistify: Au- tomating DNN model porting for on-device inference at the edge
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abdec490-9ec9-40c3-a893-4c638a35bc0f · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge A reconfigurable floating- point compute-in-memory with analog exponent pre- processes.IEEE Solid-State Circuits Letters, 2024
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7239543f-7a02-46bb-af68-8a51fc5d189c · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge A 28nm 314.6 tl- fops/w reconfigurable floating-point analog compute- in-memory macro with exponent approximation and two-stage sharing td-adc
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a24dfbc-92cc-4df8-b7c9-58ab85a82f62 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Channel pruning for accelerating very deep neural networks
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b98e9c1-f26a-465a-a724-b3405b37e79f · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Measuring massive multitask language under- standing.Proceedings of the International Conference on Learning Representations (ICLR), 2021
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b1be03e-9e80-478b-a717-7833c3102880 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Variation- aware dynamic voltage/frequency scaling
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3076f9ce-ad6b-40d4-a7f6-57c4a6c43403 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Enright Jerger
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be61a365-d8c8-4190-8266-e7c23b044d80 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Long short- term memory.Neural Comput., 9(8):1735–1780, 1997
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2da47bce-206d-4cbc-bf0d-15b02dbbbcda · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge The curious case of neural text degener- ation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 769eb812-4ae8-4723-981d-74af0549f181 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge An all-digital phase-locked loop (adpll)-based clock re- covery circuit.IEEE Journal of Solid-State Circuits, 34(8):1063–1073, 1999
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e8d0a00-0354-4f0b-b5f2-80aecb7af26f · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge LoRA: Low-Rank Adaptation of Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e817ff94-a586-40ec-b214-43b09edf3e5e · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge CAMA: energy and memory efficient automata pro- cessing in content-addressable memories
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ad978e9-5856-4fc0-bbdc-2241e8622759 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge New solutions on llm acceleration, optimization, and application
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4046c736-b798-4d42-bbf3-8fc7fd4dfbd2 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Quantized neural networks: Training neural networks with low precision weights and activations.Journal of Machine Learning Research, 18(187):1–30, 2018
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e44aefa9-87c3-44f0-9c7d-885b1066a2ea · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Transformers documentation, 2024
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd6a8791-f9bb-4329-98a7-ca39f72f9408 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Just-in-time quan- tization with processing-in-memory for efficient ml training, 2023
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62b91761-3cec-42f8-afd2-f5ed8fe933ca · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Mem: Your ai-powered assistant, 2024
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ae7c0b3-7caf-4ebf-8247-e6c83b4ee226 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Jacobs, Michael I
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4277b1f4-5a03-4635-a0e8-8f9341c6abf5 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge LLM Performance Predictors are good initializers for Architecture Search
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ba7c83e-72ad-4c44-a7ee-3a26fcd32dd4 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Robot control using llama: Bridging ai and robotics
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b1b9f69-d1c5-48c1-91d7-7e49efbd5bae · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Reinforcement learning: A survey
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6b9963e-b9fd-4107-9ffc-84269b40d97d · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Scaling Laws for Neural Language Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f2b5ecd-04d7-45cd-92e9-9f66ec8858a7 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge The breakthrough memory solutions for improved performance on LLM inference.IEEE Micro, 44(3):40–48, 2024
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ee029fc-4609-43f4-acf9-f06ae3e5facb · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Unresolved cited work
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bf5eddd-c825-4b65-be80-2450cca854f1 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Enhancing energy efficiency of multimedia applications in heterogeneous mobile multi-core pro- cessors.IEEE Trans
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ee46a44-a1a5-448b-a0a4-d3304c65b046 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge A novel gpu power model for accurate smartphone power break- down.ETRI journal, 37(1):157–164, 2015
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86b60bfc-362f-4f65-8a48-0116df482eed · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge A survey on recent os-level energy management tech- niques for mobile processing units.IEEE Trans
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b39b6851-d81f-407a-9acb-e2da44f62b94 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Autoscale: Energy efficiency optimization for stochastic edge in- ference using reinforcement learning
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83517662-138f-4b16-a548-541351317bda · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge InProceedings of the Fourteenth EuroSys Conference 2019, pages 1–15, 2019
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 408bd45e-875e-4fbf-9849-faeb181c09ed · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Distillm: Towards streamlined distillation for large language models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef5cbf70-b35e-42c5-a755-0304a781186f · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Resnet 50.Convolu- tional neural networks with swift for tensorflow: image recognition and dataset categorization, pages 63–72, 2021
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f794a6a5-a485-4ecc-afbf-392b83569178 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Automatic domain-specific soc design for autonomous unmanned aerial vehicles
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd950b1a-20b4-48bc-8374-9fbc7b8b3643 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Efficient memory management for large language model serving with pagedattention
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed56dd15-8d16-4366-9676-5900aa6c51d3 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Deep learning.nature, 521(7553):436–444, 2015
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69cebccd-b948-40af-afd5-fdc2c4638e23 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge On-chip memory technology design space explorations for mobile deep neural network ac- celerators
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e1ddf1c-7d40-4158-a804-d0e42df8fa37 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Energydx: Diag- nosing energy anomaly in mobile apps by identifying the manifestation point
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf6dc3b1-2e9c-4e64-b001-e5d0b9c9fdb9 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Deep Reinforcement Learning: An Overview
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d60f56d-b849-4140-a9e7-34fe8f407d3e · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge From system 1 to system 2: A survey of reasoning large language models, 2025
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bf5fe12-fd98-4833-86a1-03a882af4dd2 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge BAT: Behavior-Aware Human-Like Trajectory Prediction for Autonomous Driving
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1ebb0ba8-a885-4001-b269-92f25beac8d0 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Unresolved cited work
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4280d027-617c-4a5a-bd68-8ef290974005 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Par- rot: Efficient serving of llm-based applications with semantic variable
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8415a56a-8d5e-4c2f-ab47-0708dbf0f757 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd8ff174-7f99-4c6e-ae28-94630d0ee829 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge A rein- forcement learning-based power management frame- work for green computing data centers
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 059d0a55-a48b-492c-88cb-8aece664d9b8 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge A weak puf-assisted strong PUF with inherent immunity to modeling attacks and ultra-low BER.IEEE Trans
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40c8c989-53c6-4d68-8751-493ecdc5f901 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing.ACM Comput
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc9831d2-69bb-420a-bb53-fa0af3891423 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Optimizing LLM Queries in Relational Data Analytics Workloads
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98c13187-fbe6-45f1-a899-603fbd6808fd · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e8bf18e-2803-44d9-815e-950769ad6970 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge S2ta: Exploiting structured sparsity for energy-efficient mobile cnn acceleration
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc06d880-bfde-4b77-a791-4d4d3dfe2413 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Deja vu: Contextual sparsity for efficient llms at inference time
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1de46390-2a91-4b0f-b3cf-42d3ec200c22 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Prediction-guided performance-energy trade-off for in- teractive applications
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7497c9fe-741d-4040-8bf4-82b39c4554a3 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c914c224-035c-4e1c-aa41-954e5254ae86 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79bf636c-b388-40a6-b286-26ac92fe6ea4 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Pointer sentinel mixture models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9c2c25e-8a62-44d0-b1a0-e89445234836 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Gemma: Open Models Based on Gemini Research and Technology
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83756a73-5010-41ed-8ba1-e7deb2c5c203 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Your Everyday AI Companion | Microsoft Bing.https://www.bing.com/new
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3d61d069-39b8-4c12-9e0e-9c70e9469066 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Azure cognitive services - text analytics: Smart reply, 2024
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9c879a7c-19b0-46e7-b8b8-44270946c878 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Deepspeed: Advancing the science of ai through efficient training of large models, 2024
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2072150f-84ce-4f0b-a638-25af384b7116 · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Can a suit of armor conduct electricity? A new dataset for open book question answering
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c896b732-803f-40d2-8d70-e7656d596dcb · outbound
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Human-level control through deep reinforcement learning.nature, 518(7540):529– 533, 2015
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 175381e9-f139-44de-8324-94595eb56e89 · inbound
ELEVATE: Designing Human-Centered GenAI Virtual Tutors for Scalable and Inclusive Education CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.