Pith. sign in

Paper Citation Record · LEDGER

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge

As of 23 August 2026, this Paper Citation Record lists 100 of 162 outbound references and 1 inbound Pith citation observation for arXiv:2506.02847.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02847 v1

Coverage vector

measured 100 of 162 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:20:30.299592Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T07:45:22.445395Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T07:45:29.205817Z

Reference resolution

100 of 162 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved94
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3cc23a94-ad17-4c6c-92a8-25b94e6a0924 · outbound

This paper cites Rewind: A better way to search your digital life, 2024.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Rewind: A better way to search your digital life, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:29.361870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:29.361870Z digest=sha256:ab7df4ccbc7f74ba5977563db4b21e9860efde3b232c53d5affd485be76eed47

Observation 897b39d5-ddef-4220-98f7-a17bea375f7f · outbound

This paper cites An open-source framework for autonomous soc design with analog block generation.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge An open-source framework for autonomous soc design with analog block generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:29.442444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:29.442444Z digest=sha256:a6cdf5cc7fdbc23f37e30fc29cf3de66fd6a9b8ab13f3330362f10d85e0b62d0

Observation a3fd116f-ad1c-453a-82d5-acdf9a3eced5 · outbound

This paper cites The illustrated gpt-2 (visualizing trans- former language models).Jalammar.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge The illustrated gpt-2 (visualizing trans- former language models).Jalammar

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:29.637143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:29.637143Z digest=sha256:33979138d48f1388bf7309c1ba8984e7dbda29099062771a839049662360ebf8

Observation a21b22ec-6027-4945-98ae-68596982659f · outbound

This paper cites LLM in a flash: Efficient Large Language Model Inference with Limited Memory.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:29.786243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:29.786243Z digest=sha256:f1a422519d28137c9702d2e6420590e7236fc463011605ac3f36773ff62361cd

Observation 292e48e6-24ab-42f0-8ff2-9c9d8af50a64 · outbound

This paper cites Claude: A family of ai models, 2024.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Claude: A family of ai models, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:29.817687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:29.817687Z digest=sha256:16ef66b7fc0859ea709d3bc5e8766431ed6fe0048a1250ce2f56c77e2e9c78ba

Observation 64d336f0-d870-4a8c-bd09-1d635053d3fa · outbound

This paper cites Apple intelligence.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Apple intelligence

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:29.972873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:29.972873Z digest=sha256:f65e3b1d50d84dae01dea40c860f8f242b72d9e25921005f1f88a27fff0bd63e

Observation ff3fe49a-584b-442e-8c18-875471b3ad65 · outbound

This paper cites Apple siri: Virtual assistant, 2024.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Apple siri: Virtual assistant, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:29.980912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:29.980912Z digest=sha256:935de9e838defd3ddcf70bedd4e108aff4d9aa95bc181aa50d03bc2a2e267ed3

Observation 3c47924d-2585-4a07-af14-b99b5b35ee21 · outbound

This paper cites SliceGPT: Compress Large Language Models by Deleting Rows and Columns.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:29.984809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:29.984809Z digest=sha256:50dddb710eddc0351ff07e66990981a70d6c1d84a150c0e5919accb5b3b6bfc4

Observation 37fd1495-cc8f-4f20-9b83-914429bddfe3 · outbound

This paper cites 25.1 a fully synthesizable distributed and scalable all-digital ldo in 10nm cmos.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge 25.1 a fully synthesizable distributed and scalable all-digital ldo in 10nm cmos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:29.988251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:29.988251Z digest=sha256:6ccdcaa3582d8f6740dec18f44bac8f2cb05975348a21174ee52a089fba1d8ee

Observation e7146166-892c-4f84-bf31-4346bce7de12 · outbound

This paper cites {NeuOS}: A {Latency- Predictable}{Multi-Dimensional} optimization frame- work for {DNN-driven} autonomous systems.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge {NeuOS}: A {Latency- Predictable}{Multi-Dimensional} optimization frame- work for {DNN-driven} autonomous systems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.033790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.033790Z digest=sha256:c93314f9a58d2ddd1a0af1c58bd53e9cc26ec83436b1cdb76222571099427369

Observation 18e0b18a-e656-4515-b77f-133a1022c919 · outbound

This paper cites Bender, Timnit Gebru, Angelina McMillan- Major, and Shmargaret Shmitchell.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Bender, Timnit Gebru, Angelina McMillan- Major, and Shmargaret Shmitchell

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.037121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.037121Z digest=sha256:452a6a0897f468acb7ff424565290264dc85a38488dfcf701a1e5d3eec26a5c5

Observation 095e4547-b476-4382-8ec6-99492013f2b9 · outbound

This paper cites PIQA: reasoning about physical commonsense in natural language.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge PIQA: reasoning about physical commonsense in natural language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.040052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.040052Z digest=sha256:fd42de9013f798e9831ad5bab4eaa961b84d8eaef75dad56a35dd6c3229881a4

Observation 0ceb2af1-898b-4870-97cc-203e0d496c8a · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge On the Opportunities and Risks of Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.043082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.043082Z digest=sha256:3c38011e1417dd37a69a7cce7b429a543b55e20a305ea89513e26677a9df4b97

Observation da7612c2-af60-4652-a382-cd1498fb8143 · outbound

This paper cites An estimate of an upper bound for the entropy of english.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge An estimate of an upper bound for the entropy of english

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.046628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.046628Z digest=sha256:04c5157d59e7b9fc72d2343ced1d7bfe89c1b75ecffe6700bbccc945338798ee

Observation 7f5a3be4-0875-4749-9f67-3d1821e7fdb4 · outbound

This paper cites an unresolved cited work.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.049551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.049551Z digest=sha256:e019fe222d644a137004b4d7c75b6c5a3e6e859a0873b8b38b4a27a256f65f09

Observation 6e89f9a1-9fd3-4cf6-a3df-6948022d7d73 · outbound

This paper cites A fre- quency compensation scheme for ldo voltage regula- tors.IEEE Transactions on Circuits and Systems I: Regular Papers, 51(6):1041–1050, 2004.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge A fre- quency compensation scheme for ldo voltage regula- tors.IEEE Transactions on Circuits and Systems I: Regular Papers, 51(6):1041–1050, 2004

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.052735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.052735Z digest=sha256:839efc5ef214e934fa603dbd94f483c145af50f85af1ad70024d2d733257bd00

Observation 3838fe1c-f4e0-4005-9923-c6954dca22e9 · outbound

This paper cites Bge m3-embedding: Multi- lingual, multi-functionality, multi-granularity text em- beddings through self-knowledge distillation, 2024.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Bge m3-embedding: Multi- lingual, multi-functionality, multi-granularity text em- beddings through self-knowledge distillation, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.055847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.055847Z digest=sha256:b5192bf862b3e8f740fa1169845442519296e3fd2ab5596ab019c94b1efc7566

Observation a3ea6e24-12bb-45ea-8dfd-9c691b9bacf5 · outbound

This paper cites Op- timizing energy efficiency of browsers in energy-aware scheduling-enabled mobile devices.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Op- timizing energy efficiency of browsers in energy-aware scheduling-enabled mobile devices

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.058717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.058717Z digest=sha256:6fb85d02388330ddae37a576b5ee91f687e0dd5dbfd7560da0fa644c84ced095

Observation 872e40fe-bb1c-4730-8535-aacdb0088def · outbound

This paper cites an unresolved cited work.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.061530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.061530Z digest=sha256:89245e85d1b5f1940f0c584eb863f9b89dc7889902ea97eec56ce6493f41925d

Observation 77a962fe-6853-4879-b4d7-000581fb08d1 · outbound

This paper cites ML.ENERGY leaderboard.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge ML.ENERGY leaderboard

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.064298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.064298Z digest=sha256:4bd26e3b166a79407237a41d77f0ed62a0dda3e6074ede1c58a454ec8ba4d7a9

Observation 45893bdb-be31-49d0-9e61-c067e8026f8c · outbound

This paper cites Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche, Eliza Ruther- ford, Tom Hennigan, Matthew J.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche, Eliza Ruther- ford, Tom Hennigan, Matthew J

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.066924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.066924Z digest=sha256:9c1e34a9060816c99141eb48a0d05c1f568133eb330f8b699102b693d4e141f5

Observation 7fc02c78-a21c-4b6f-aad3-4c110cdd1624 · outbound

This paper cites Boolq: Exploring the surprising difficulty of natural yes/no questions.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Boolq: Exploring the surprising difficulty of natural yes/no questions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.069780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.069780Z digest=sha256:1a3f02e1a1bfdbe3deb5e742b1053bf50c929589432bcf570d420c012cf354c4

Observation 64ffa5f7-1128-4cae-b93d-847c1d46ee0a · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.072546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.072546Z digest=sha256:3240fc979a4f5871207eab33764ef50d31c6f0ac109c67d25910ad4de66bd039

Observation d4b4ae27-2128-4ce2-a496-5b21c0ba3f82 · outbound

This paper cites Track and reduce co2 emissions from your computing.https://codecarbon.io/, 2023.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Track and reduce co2 emissions from your computing.https://codecarbon.io/, 2023

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.075288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.075288Z digest=sha256:21f3506909e61671879bda0d8023861d7515174f4dbb60311a84b36fc10f1b83

Observation 796f6d98-7a46-477d-a3bb-58d44d4844a3 · outbound

This paper cites Llama 2: Inferencing on a single gpu.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Llama 2: Inferencing on a single gpu

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.077959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.077959Z digest=sha256:372f5c14a06d95e2e655d985821c60aa193a8c65d98bc93a322d613eea3c427f

Observation 604a50a1-60b3-4f03-9137-e059d4a4020c · outbound

This paper cites Llm.int8(): 8-bit matrix multiplication for transformers at scale.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Llm.int8(): 8-bit matrix multiplication for transformers at scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.081219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.081219Z digest=sha256:ccd66cace550c2c240799c48444f402dc1c8b550f2ba122d603aaf46641b6bbf

Observation 64368e54-f293-4069-ae4d-d84d825cc5ef · outbound

This paper cites Pruner- zero: Evolving symbolic pruning metric from scratch for large language models.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Pruner- zero: Evolving symbolic pruning metric from scratch for large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.083937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.083937Z digest=sha256:691d45d1b768b53e3b37497be3d2aac0fa9c7732a1ea8c3dbbe5e62f9ba7d4a9

Observation 8f2949f7-6d06-4034-a4f3-653ce46f9761 · outbound

This paper cites User-aware frame rate man- agement in android smartphones.ACM Trans.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge User-aware frame rate man- agement in android smartphones.ACM Trans

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.086587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.086587Z digest=sha256:a95be5eebb3fa674a014fbba854fd31863d9d5b935fcdfa8fcc9b6e838f89a54

Observation 041a2a81-ab56-44b6-b9db-db536e6254ae · outbound

This paper cites Bradley Chen, and Margo I.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Bradley Chen, and Margo I

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.089369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.089369Z digest=sha256:226f90c389c3614a571c535718c5327ccbcd91dce0452ef4c697a997c75908bb

Observation 9583e905-eaa1-44e8-9835-575700d5f855 · outbound

This paper cites Depgraph: Towards any struc- tural pruning.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Depgraph: Towards any struc- tural pruning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.092261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.092261Z digest=sha256:f01132434ce18aeb43ebdff863dc943147e24e508ce533e0356828e095b4a594

Observation de9f1fce-d890-4ca5-b48d-1f823c8af701 · outbound

This paper cites Gpt-3: Its nature, scope, limits, and consequences.Minds and Machines, 30:681–694, 2020.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Gpt-3: Its nature, scope, limits, and consequences.Minds and Machines, 30:681–694, 2020

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.095081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.095081Z digest=sha256:a64f4958f71766f49fa02508c3e1d2bc1578128dd88a688edc183f2acd2c3485

Observation fb189c35-5f9d-45bd-9114-056fddd214de · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one- shot.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Sparsegpt: Massive language models can be accurately pruned in one- shot

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.097575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.097575Z digest=sha256:eb1e7e3c9c78e5e1d117d48d1f53d069dd27a41024ee026fe18fcea9207b6a43

Observation b55ce9d8-21c3-4720-9cd7-b87bd6aeaafd · outbound

This paper cites Beam search strategies for neural machine translation.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Beam search strategies for neural machine translation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.100144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.100144Z digest=sha256:7b007d7c187f46f55d9bdd6d7a2f8df411aa93007e24f044e358ebff8adedf8a

Observation 69364cba-1976-45da-a39a-b3df3fead741 · outbound

This paper cites Drive like a human: Rethinking autonomous driving with large language models.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Drive like a human: Rethinking autonomous driving with large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.103379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.103379Z digest=sha256:f1a061e75cf89959f300b73a03e0f394e3444d78da207617caf099281ea2f0f6

Observation c860af55-87db-4946-83fd-00f015302dff · outbound

This paper cites Openllama: An open reproduction of llama.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Openllama: An open reproduction of llama

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.106009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.106009Z digest=sha256:39be838bcc5229ae1ca46db97fe3594c09b5d197380fa7159d1904c7e8c41d8b

Observation b4d77d22-77ec-4e93-8ea0-00aa98e048e2 · outbound

This paper cites Mahoney, and Kurt Keutzer.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Mahoney, and Kurt Keutzer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.108661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.108661Z digest=sha256:28f9e2f6f12b63723b3b1283bd9085a9543d6cc4e703b536a3615500976f78c7

Observation f17d1f1f-d07a-4a6d-8309-55bc9119823b · outbound

This paper cites GitHub Copilot: Your AI pair programmer.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge GitHub Copilot: Your AI pair programmer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.111229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.111229Z digest=sha256:a80dcd7b598f814f2cc0f7ebc7ccc934d5362743e2690247d961d3855c0c4e47

Observation f64bbe97-1e6f-42d7-84fa-8450a8f60850 · outbound

This paper cites Goodfellow, Yoshua Bengio, and Aaron C.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Goodfellow, Yoshua Bengio, and Aaron C

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.114160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.114160Z digest=sha256:f22eb45a4d132523a115e8976ff5375becb407cbac4df69345c2525cc6ec8bc2

Observation 5b3382f0-4364-4c1c-9d32-74e10676352e · outbound

This paper cites Google assistant, 2024.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Google assistant, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.117045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.117045Z digest=sha256:3ff9a8f6a530cda481c969f4a2f50583a23a339e6830ceb247b10ac6d78257ec

Observation 1f2bac2c-c034-4d83-a651-89bceb96265b · outbound

This paper cites Ml kit smart reply, 2024.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Ml kit smart reply, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.119676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.119676Z digest=sha256:465e4bc4ad964bed23015f2793c36b8b68fbb83b95c0789e65d320ee21534d86

Observation 1cf2c75c-bcf0-4cef-ae89-6262d90269b6 · outbound

This paper cites Minillm: Knowledge distillation of large language models.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Minillm: Knowledge distillation of large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.122280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.122280Z digest=sha256:799bf2bd3ff665e537b3f7189b9024cf47c1d14ef3eb9a4bad6610c3b1f42868

Observation 48647cfb-269f-4b74-8277-6b6b16bd0223 · outbound

This paper cites Mistify: Au- tomating DNN model porting for on-device inference at the edge.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Mistify: Au- tomating DNN model porting for on-device inference at the edge

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.125464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.125464Z digest=sha256:56a16fec63e62ebd2f9b20ed924a9b173c25f9ba5e31361af9a9fcc46041f055

Observation abdec490-9ec9-40c3-a893-4c638a35bc0f · outbound

This paper cites A reconfigurable floating- point compute-in-memory with analog exponent pre- processes.IEEE Solid-State Circuits Letters, 2024.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge A reconfigurable floating- point compute-in-memory with analog exponent pre- processes.IEEE Solid-State Circuits Letters, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.128612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.128612Z digest=sha256:a4289afe3d37ad2ed36faa2e4c8593bb98b8a272fa95ff0316b0282f0d7cba7c

Observation 7239543f-7a02-46bb-af68-8a51fc5d189c · outbound

This paper cites A 28nm 314.6 tl- fops/w reconfigurable floating-point analog compute- in-memory macro with exponent approximation and two-stage sharing td-adc.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge A 28nm 314.6 tl- fops/w reconfigurable floating-point analog compute- in-memory macro with exponent approximation and two-stage sharing td-adc

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.131269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.131269Z digest=sha256:1aeaf0e7a151c542eb2da36c842de8c730cf53ff48a918d1b762ad9abd0d4445

Observation 7a24dfbc-92cc-4df8-b7c9-58ab85a82f62 · outbound

This paper cites Channel pruning for accelerating very deep neural networks.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Channel pruning for accelerating very deep neural networks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.133896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.133896Z digest=sha256:92c4c568861866a6ca9a6c317b822a32f38060c099f137a9e28088565a7220e0

Observation 3b98e9c1-f26a-465a-a724-b3405b37e79f · outbound

This paper cites Measuring massive multitask language under- standing.Proceedings of the International Conference on Learning Representations (ICLR), 2021.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Measuring massive multitask language under- standing.Proceedings of the International Conference on Learning Representations (ICLR), 2021

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.136923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.136923Z digest=sha256:03a0b33faf10046108d4f56bf1414f96d89ac0ef8dc7f56c796530e52b401aea

Observation 2b1be03e-9e80-478b-a717-7833c3102880 · outbound

This paper cites Variation- aware dynamic voltage/frequency scaling.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Variation- aware dynamic voltage/frequency scaling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.139632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.139632Z digest=sha256:f970d6b300b01b90f3022da836065cda62f6dc0789b66fd76397719120c09f69

Observation 3076f9ce-ad6b-40d4-a7f6-57c4a6c43403 · outbound

This paper cites Enright Jerger.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Enright Jerger

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.144325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.144325Z digest=sha256:0cd37c2e980c21b72b1e31bb842a8410de275290e0e637d8021811a3ae30a8dc

Observation be61a365-d8c8-4190-8266-e7c23b044d80 · outbound

This paper cites Long short- term memory.Neural Comput., 9(8):1735–1780, 1997.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Long short- term memory.Neural Comput., 9(8):1735–1780, 1997

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.147171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.147171Z digest=sha256:1322dd34fd362d381a6b28089f9a0644d1f3f7160ab31fb009542f97d811d416

Observation 2da47bce-206d-4cbc-bf0d-15b02dbbbcda · outbound

This paper cites The curious case of neural text degener- ation.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge The curious case of neural text degener- ation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.150467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.150467Z digest=sha256:22c9cf2b3d947733a117691c1dd322c062826e461cd16718e315ddb24b59bc14

Observation 769eb812-4ae8-4723-981d-74af0549f181 · outbound

This paper cites An all-digital phase-locked loop (adpll)-based clock re- covery circuit.IEEE Journal of Solid-State Circuits, 34(8):1063–1073, 1999.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge An all-digital phase-locked loop (adpll)-based clock re- covery circuit.IEEE Journal of Solid-State Circuits, 34(8):1063–1073, 1999

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.153381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.153381Z digest=sha256:1ccac546b4f9ae1ce1631c32483bf28657e690a03a9844125ed6531c3e954ed4

Observation 0e8d0a00-0354-4f0b-b5f2-80aecb7af26f · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge LoRA: Low-Rank Adaptation of Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.156139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.156139Z digest=sha256:07fc0b51fae5eb9618098a70f8a8177c7f1c0cc52eeda26689abba65e15406d8

Observation e817ff94-a586-40ec-b214-43b09edf3e5e · outbound

This paper cites CAMA: energy and memory efficient automata pro- cessing in content-addressable memories.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge CAMA: energy and memory efficient automata pro- cessing in content-addressable memories

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.159598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.159598Z digest=sha256:e34b61be101942bbac69b82a7ae8c728548c86fe20a87f1fa21c5f8dd7b58448

Observation 7ad978e9-5856-4fc0-bbdc-2241e8622759 · outbound

This paper cites New solutions on llm acceleration, optimization, and application.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge New solutions on llm acceleration, optimization, and application

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.162218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.162218Z digest=sha256:92ebe98c51b4eb12c5d3f2bd2dda6a979e4d1e3218709ec00496ae10943c0719

Observation 4046c736-b798-4d42-bbf3-8fc7fd4dfbd2 · outbound

This paper cites Quantized neural networks: Training neural networks with low precision weights and activations.Journal of Machine Learning Research, 18(187):1–30, 2018.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Quantized neural networks: Training neural networks with low precision weights and activations.Journal of Machine Learning Research, 18(187):1–30, 2018

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.164863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.164863Z digest=sha256:2a1dde80d36fe9369ba96f7f5739a8dedc3aba06ec4800c748844da82b73b34e

Observation e44aefa9-87c3-44f0-9c7d-885b1066a2ea · outbound

This paper cites Transformers documentation, 2024.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Transformers documentation, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.167284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.167284Z digest=sha256:5778a0228ccb797573ae1908f934d9c6b4691d7b56fe556e7e9e53df51b5e86b

Observation dd6a8791-f9bb-4329-98a7-ca39f72f9408 · outbound

This paper cites Just-in-time quan- tization with processing-in-memory for efficient ml training, 2023.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Just-in-time quan- tization with processing-in-memory for efficient ml training, 2023

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.170049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.170049Z digest=sha256:66b504cd55156eb958ff3d5ca8eea82ab09452398f04d50feff3900e3cb8a5e6

Observation 62b91761-3cec-42f8-afd2-f5ed8fe933ca · outbound

This paper cites Mem: Your ai-powered assistant, 2024.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Mem: Your ai-powered assistant, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.172888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.172888Z digest=sha256:0b02b835205ec9dfb9618cfe32c0f5914760a4c65b33da1236b03d5fb72a74f9

Observation 0ae7c0b3-7caf-4ebf-8247-e6c83b4ee226 · outbound

This paper cites Jacobs, Michael I.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Jacobs, Michael I

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.175585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.175585Z digest=sha256:22913ff72fa09d0fe78c8c2e6894720c9ab194cfd5d365d1734c389de312616a

Observation 4277b1f4-5a03-4635-a0e8-8f9341c6abf5 · outbound

This paper cites LLM Performance Predictors are good initializers for Architecture Search.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge LLM Performance Predictors are good initializers for Architecture Search

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.178185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.178185Z digest=sha256:a813d21ddb5bc13b8634a1aa42988d6a30e340d0df2df29e18ee8c288bd87958

Observation 6ba7c83e-72ad-4c44-a7ee-3a26fcd32dd4 · outbound

This paper cites Robot control using llama: Bridging ai and robotics.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Robot control using llama: Bridging ai and robotics

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.181453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.181453Z digest=sha256:69f803574c563f5b01e9784ca96992c88b3151089952a5448278feb3966ae864

Observation 0b1b9f69-d1c5-48c1-91d7-7e49efbd5bae · outbound

This paper cites Reinforcement learning: A survey.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Reinforcement learning: A survey

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.184097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.184097Z digest=sha256:a4b0d0717ccab9372f7adc1f3cd2a81c8a050644fce1fc34216a5a892a72e984

Observation a6b9963e-b9fd-4107-9ffc-84269b40d97d · outbound

This paper cites Scaling Laws for Neural Language Models.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Scaling Laws for Neural Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.186985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.186985Z digest=sha256:fed5482d29dedcbef36ff5da049e827106a3a340bd049c5b7c372d9c6a612c84

Observation 7f2b5ecd-04d7-45cd-92e9-9f66ec8858a7 · outbound

This paper cites The breakthrough memory solutions for improved performance on LLM inference.IEEE Micro, 44(3):40–48, 2024.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge The breakthrough memory solutions for improved performance on LLM inference.IEEE Micro, 44(3):40–48, 2024

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.189680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.189680Z digest=sha256:742766764001b843227152bbc38e1f6e281cb6aaf39b44e91315320981245637

Observation 3ee029fc-4609-43f4-acf9-f06ae3e5facb · outbound

This paper cites an unresolved cited work.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.192366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.192366Z digest=sha256:9222dd765d1363f13ea35a8d926b0d96cbfc606b22cfff2c28876ea4cd7f512d

Observation 2bf5eddd-c825-4b65-be80-2450cca854f1 · outbound

This paper cites Enhancing energy efficiency of multimedia applications in heterogeneous mobile multi-core pro- cessors.IEEE Trans.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Enhancing energy efficiency of multimedia applications in heterogeneous mobile multi-core pro- cessors.IEEE Trans

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.195212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.195212Z digest=sha256:b8fc0e47156e04d4bf57ac0421a22de51982dcda9be87cd4649d90c7997f7ce6

Observation 9ee46a44-a1a5-448b-a0a4-d3304c65b046 · outbound

This paper cites A novel gpu power model for accurate smartphone power break- down.ETRI journal, 37(1):157–164, 2015.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge A novel gpu power model for accurate smartphone power break- down.ETRI journal, 37(1):157–164, 2015

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.198205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.198205Z digest=sha256:64534407a03fb3441cc0c434ecb2e99847eb74211ab18454c5f869ad89b5aa22

Observation 86b60bfc-362f-4f65-8a48-0116df482eed · outbound

This paper cites A survey on recent os-level energy management tech- niques for mobile processing units.IEEE Trans.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge A survey on recent os-level energy management tech- niques for mobile processing units.IEEE Trans

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.200853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.200853Z digest=sha256:0997d04b2738dda12b9b55ada496db2a30fd61b181e601acf6006c2e9b734049

Observation b39b6851-d81f-407a-9acb-e2da44f62b94 · outbound

This paper cites Autoscale: Energy efficiency optimization for stochastic edge in- ference using reinforcement learning.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Autoscale: Energy efficiency optimization for stochastic edge in- ference using reinforcement learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.203498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.203498Z digest=sha256:44abde7b8abdcd601628db1cb4f85d4709da6cf669a9d51fafefd57a53f85ed7

Observation 83517662-138f-4b16-a548-541351317bda · outbound

This paper cites InProceedings of the Fourteenth EuroSys Conference 2019, pages 1–15, 2019.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge InProceedings of the Fourteenth EuroSys Conference 2019, pages 1–15, 2019

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.206434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.206434Z digest=sha256:8af00225a5fd56505e47be837bf4ae718c1ec19b9a10ef1bfa0401f0fb2436ff

Observation 408bd45e-875e-4fbf-9849-faeb181c09ed · outbound

This paper cites Distillm: Towards streamlined distillation for large language models.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Distillm: Towards streamlined distillation for large language models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.209192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.209192Z digest=sha256:5800403ccf9ffcfe7b9962ba64b6b6ce867e1d4b8f41cfc6c54f02b9b2aba7f3

Observation ef5cbf70-b35e-42c5-a755-0304a781186f · outbound

This paper cites Resnet 50.Convolu- tional neural networks with swift for tensorflow: image recognition and dataset categorization, pages 63–72, 2021.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Resnet 50.Convolu- tional neural networks with swift for tensorflow: image recognition and dataset categorization, pages 63–72, 2021

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.212076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.212076Z digest=sha256:54d36fb9cc335e783ba7519db3474dd255f4cf8ce8c1ac57de69f4399940e964

Observation f794a6a5-a485-4ecc-afbf-392b83569178 · outbound

This paper cites Automatic domain-specific soc design for autonomous unmanned aerial vehicles.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Automatic domain-specific soc design for autonomous unmanned aerial vehicles

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.215564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.215564Z digest=sha256:f926f9f21549c603b820b47276dfa3efe656f2195ee20dda48d259f95e07dedf

Observation bd950b1a-20b4-48bc-8374-9fbc7b8b3643 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Efficient memory management for large language model serving with pagedattention

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.221498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.221498Z digest=sha256:1d973b0d2c00bc45eed25bd5e1bd1bbd867306a804a08b76ee917d627076106c

Observation ed56dd15-8d16-4366-9676-5900aa6c51d3 · outbound

This paper cites Deep learning.nature, 521(7553):436–444, 2015.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Deep learning.nature, 521(7553):436–444, 2015

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.224331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.224331Z digest=sha256:47486eaacd174b474994c212ae41612b1a65187794ad86332772a7714df067b5

Observation 69cebccd-b948-40af-afd5-fdc2c4638e23 · outbound

This paper cites On-chip memory technology design space explorations for mobile deep neural network ac- celerators.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge On-chip memory technology design space explorations for mobile deep neural network ac- celerators

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.227272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.227272Z digest=sha256:66c961e318d081747082fb9930bccc98e73ddb73f471c8c42cc6f159ce750b5a

Observation 8e1ddf1c-7d40-4158-a804-d0e42df8fa37 · outbound

This paper cites Energydx: Diag- nosing energy anomaly in mobile apps by identifying the manifestation point.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Energydx: Diag- nosing energy anomaly in mobile apps by identifying the manifestation point

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.230317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.230317Z digest=sha256:6436c116feb302ff75e40c7ae73a6e0ded86e320f6df6268ee3d2a425e5333e5

Observation bf6dc3b1-2e9c-4e64-b001-e5d0b9c9fdb9 · outbound

This paper cites Deep Reinforcement Learning: An Overview.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Deep Reinforcement Learning: An Overview

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.233137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.233137Z digest=sha256:9fa57729c07318e661c1eadc97720c26c57c760161719df442bd8d2133904ed8

Observation 7d60f56d-b849-4140-a9e7-34fe8f407d3e · outbound

This paper cites From system 1 to system 2: A survey of reasoning large language models, 2025.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge From system 1 to system 2: A survey of reasoning large language models, 2025

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.236152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.236152Z digest=sha256:6a56f0ed5fe52ecfe9ac4a46c3bb95ac4a386778a4d8f31736cf49678981b24f

Observation 1bf5fe12-fd98-4833-86a1-03a882af4dd2 · outbound

This paper cites BAT: Behavior-Aware Human-Like Trajectory Prediction for Autonomous Driving.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge BAT: Behavior-Aware Human-Like Trajectory Prediction for Autonomous Driving

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:20:30.967973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:20:30.238811Z digest=sha256:8181b2ca1515ec460c3dd83fca2c3ad81bf0f6b6ae662cdb480c6d2fb0b88b84

Observation 1ebb0ba8-a885-4001-b269-92f25beac8d0 · outbound

This paper cites an unresolved cited work.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.241803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.241803Z digest=sha256:1c7d1b898757868baaaf914973fee2a6c639b43b486a1e14f0ce420817fed2d2

Observation 4280d027-617c-4a5a-bd68-8ef290974005 · outbound

This paper cites Par- rot: Efficient serving of llm-based applications with semantic variable.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Par- rot: Efficient serving of llm-based applications with semantic variable

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.244455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.244455Z digest=sha256:0460989f8fab1e3c359c112c2cda9e0d00367f42c6e0c7faf680d55be89f58d1

Observation 8415a56a-8d5e-4c2f-ab47-0708dbf0f757 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.250265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.250265Z digest=sha256:5fdfc344b95a8e93fd2fb8ab0979ddf1f4e651a4303edbc2582e93ebfea4e5ed

Observation fd8ff174-7f99-4c6e-ae28-94630d0ee829 · outbound

This paper cites A rein- forcement learning-based power management frame- work for green computing data centers.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge A rein- forcement learning-based power management frame- work for green computing data centers

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.253352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.253352Z digest=sha256:2c6700825af68cf8c7de9481695855fd3e95f774620052e35e1a2e1816da0469

Observation 059d0a55-a48b-492c-88cb-8aece664d9b8 · outbound

This paper cites A weak puf-assisted strong PUF with inherent immunity to modeling attacks and ultra-low BER.IEEE Trans.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge A weak puf-assisted strong PUF with inherent immunity to modeling attacks and ultra-low BER.IEEE Trans

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.256232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.256232Z digest=sha256:10195a4931829a76388eebbf2e3238dcfe74626f96735963a2c58701d0fb5d08

Observation 40c8c989-53c6-4d68-8751-493ecdc5f901 · outbound

This paper cites Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing.ACM Comput.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing.ACM Comput

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.259016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.259016Z digest=sha256:a89e01a693314d7f6c8a04da44357b174e243f503b90a2c2d10e50a29aebc07b

Observation dc9831d2-69bb-420a-bb53-fa0af3891423 · outbound

This paper cites Optimizing LLM Queries in Relational Data Analytics Workloads.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Optimizing LLM Queries in Relational Data Analytics Workloads

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.261857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.261857Z digest=sha256:e69d04c0c99b0c67579e1be3dcb43d71480d17021e0153cebc32c600cbdfe583

Observation 98c13187-fbe6-45f1-a899-603fbd6808fd · outbound

This paper cites P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.265155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.265155Z digest=sha256:453a5a785075517bff2e7f80ac279572695fc122239f8d7d2733cfadcc3d7600

Observation 5e8bf18e-2803-44d9-815e-950769ad6970 · outbound

This paper cites S2ta: Exploiting structured sparsity for energy-efficient mobile cnn acceleration.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge S2ta: Exploiting structured sparsity for energy-efficient mobile cnn acceleration

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.268065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.268065Z digest=sha256:b46b0af60e9ec93a18bfaf9f8449b99290cc63813bf574b691f29919f8970216

Observation cc06d880-bfde-4b77-a791-4d4d3dfe2413 · outbound

This paper cites Deja vu: Contextual sparsity for efficient llms at inference time.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Deja vu: Contextual sparsity for efficient llms at inference time

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.270805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.270805Z digest=sha256:9bb6d43cc50dd16b4e11141f9dc4a0b7c0d3880b650ac61b4eabf259facf982d

Observation 1de46390-2a91-4b0f-b3cf-42d3ec200c22 · outbound

This paper cites Prediction-guided performance-energy trade-off for in- teractive applications.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Prediction-guided performance-energy trade-off for in- teractive applications

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.273473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.273473Z digest=sha256:f2bcae481f0945c71b055bc59b4d05209dd665a1a7de4c67ea59f862bdda1a4e

Observation 7497c9fe-741d-4040-8bf4-82b39c4554a3 · outbound

This paper cites Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.276218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.276218Z digest=sha256:2db96b9b197d0c7dd008f3b2f7e84231e4acfcb8a5e086df73529c94c297fe39

Observation c914c224-035c-4e1c-aa41-954e5254ae86 · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.278863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.278863Z digest=sha256:794764dedd0305d4d06489e34d1c5642c0183a356cc445dc73ac3436047286a5

Observation 79bf636c-b388-40a6-b286-26ac92fe6ea4 · outbound

This paper cites Pointer sentinel mixture models.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Pointer sentinel mixture models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.281679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.281679Z digest=sha256:3bafd1ca722d8ca31397de18831293211c74faecd44ac62dbaa2838d1f8f4fe1

Observation b9c2c25e-8a62-44d0-b1a0-e89445234836 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Gemma: Open Models Based on Gemini Research and Technology

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:30.284503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:30.284503Z digest=sha256:dfd80342271593a006df575c469a443f0948e82d3eee942ce1bf2ccaf092a1b1

Observation 83756a73-5010-41ed-8ba1-e7deb2c5c203 · outbound

This paper cites Your Everyday AI Companion | Microsoft Bing.https://www.bing.com/new.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Your Everyday AI Companion | Microsoft Bing.https://www.bing.com/new

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:31.575245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:20:30.287740Z digest=sha256:75bb61fac10fa547e4a68711a7b8d6fcf2e5e447e864d715d17769dfedb2f624

Observation 3d61d069-39b8-4c12-9e0e-9c70e9469066 · outbound

This paper cites Azure cognitive services - text analytics: Smart reply, 2024.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Azure cognitive services - text analytics: Smart reply, 2024

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:31.564613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:20:30.290561Z digest=sha256:e6ff564a3e9d8e254d05de0de175bfc7f1f1712b4ca20ab5ac8ac4fab86fcb9e

Observation 9c879a7c-19b0-46e7-b8b8-44270946c878 · outbound

This paper cites Deepspeed: Advancing the science of ai through efficient training of large models, 2024.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Deepspeed: Advancing the science of ai through efficient training of large models, 2024

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:31.554368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:20:30.293309Z digest=sha256:f94a39f67447658da053dc9ede111ac8d4b4457d31cc343288d747aa03d1f9ea

Observation 2072150f-84ce-4f0b-a638-25af384b7116 · outbound

This paper cites Can a suit of armor conduct electricity? A new dataset for open book question answering.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Can a suit of armor conduct electricity? A new dataset for open book question answering

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:31.543370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:20:30.296304Z digest=sha256:dc5eb36694a492b68b4afb7841743052e82444090542db2f07b5d8ab415dd983

Observation c896b732-803f-40d2-8d70-e7656d596dcb · outbound

This paper cites Human-level control through deep reinforcement learning.nature, 518(7540):529– 533, 2015.

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge Human-level control through deep reinforcement learning.nature, 518(7540):529– 533, 2015

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:31.531742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:20:30.299592Z digest=sha256:7793ddf33528b0feac5cf4d8012c0eaeaf134c76da607726da942ba4448bac74

Pith citing papers

Observation 175381e9-f139-44de-8324-94595eb56e89 · inbound

ELEVATE: Designing Human-Centered GenAI Virtual Tutors for Scalable and Inclusive Education cites this paper.

ELEVATE: Designing Human-Centered GenAI Virtual Tutors for Scalable and Inclusive Education CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:45:29.207153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-01T07:45:22.445395Z digest=sha256:ba1d6e199e5988c76d15ab43cb8a1e476bb6fd5e51450082b76eda3f7b1ae449