Pith. sign in

Paper Citation Record · LEDGER

Efficient Large Language Models: A Survey

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 59 inbound Pith citation observations for arXiv:2312.03863.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.03863 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 59 of 59 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:15:17.326433Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

24
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f881ece7-0004-402a-a558-b6a8cdfae25a · inbound

A Survey on the Memory Mechanism of Large Language Model based Agents cites this paper.

A Survey on the Memory Mechanism of Large Language Model based Agents Efficient Large Language Models: A Survey

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:21:39.755184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T07:21:39.440092Z digest=sha256:86177fc35e3edaaa1f2b98ed59eb625e31d4af0ad5a6dee8cdda1507adf5ebc4

Observation c67e25e3-312e-4b6c-be31-feb3e71905fe · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Efficient Large Language Models: A Survey

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.686744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:9974ec3fbb4931582001e15f1787db0eed9b4d04633d02ec3e6350bfb4b9f726

Observation 7c916fd0-fe42-40e1-a36c-67926ae38a7d · inbound

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap cites this paper.

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap Efficient Large Language Models: A Survey

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:08:20.777408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T19:07:21.016824Z digest=sha256:fb06ca2f124d4df346e397b12a9da669c156717339de1734854e6164f0d4f921

Observation 162cc3db-c0d0-45ed-b514-07c4426e776d · inbound

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree cites this paper.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Efficient Large Language Models: A Survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.744899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.744899Z digest=sha256:c7bff3f12b71a4b276a5482e98cf8ddf4344a70c4b6fc8c719faa50703464323

Observation 40b2cb96-c13c-4201-abee-11871f7f3d25 · inbound

A Survey on Inference Optimization Techniques for Mixture of Experts Models cites this paper.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Efficient Large Language Models: A Survey

Reference 183

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:36.051126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:36.051126Z digest=sha256:99fac657e34cc794ed7b46c5c2ee3d1b66d5a278f5773b770528366fb9384b0c

Observation f367fe39-b176-4c2b-91f3-8b6285a68126 · inbound

All-in-One Tuning and Structural Pruning for Domain-Specific LLMs cites this paper.

All-in-One Tuning and Structural Pruning for Domain-Specific LLMs Efficient Large Language Models: A Survey

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T12:17:56.782840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:17:56.782840Z digest=sha256:94489241c340d9b1beb010deb65809739a2543e639f2c6a03f6cfa8995f1c882

Observation 669082bf-1ecf-42f6-8c71-412ab3137cb0 · inbound

QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning cites this paper.

QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning Efficient Large Language Models: A Survey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:22:41.334325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:22:41.334325Z digest=sha256:01d3e1d9f06591473a8e8a7f44356e21fc367a640a29a79809f964a4304df368

Observation caf0874c-a945-4ea1-ba4b-f90e7d3a6bb5 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Efficient Large Language Models: A Survey

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:46.580863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:46.580863Z digest=sha256:98ef8a26db4a19d172bb0a789e25f378bc3e11e55c4e06e6612a79903b65d9f5

Observation 6ae001d9-89f7-42cf-959a-c864939fc4cc · inbound

TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models cites this paper.

TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models Efficient Large Language Models: A Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T05:35:54.107705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:35:54.107705Z digest=sha256:bcf96701d8ee516ffc3e5ac18ed157714f1fdff3d158e77aa22831d50f161590

Observation aace18d2-f7df-4958-aab2-d7478aaebc42 · inbound

Pivoting Factorization: A Compact Meta Low-Rank Representation of Sparsity for Efficient Inference in Large Language Models cites this paper.

Pivoting Factorization: A Compact Meta Low-Rank Representation of Sparsity for Efficient Inference in Large Language Models Efficient Large Language Models: A Survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T21:31:36.265172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:31:36.265172Z digest=sha256:5dde848d2ec7557263c5450a446361e4454913a59d68cd5a26ef48aef6c37a10

Observation 1ce125b3-7d8a-4f09-82db-f6f0270329c2 · inbound

Learning Model Successors cites this paper.

Learning Model Successors Efficient Large Language Models: A Survey

Reference 168

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:59.088694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:59.088694Z digest=sha256:6dfe1a96cfe0a164aef1b9794272680e831819d650ceebbd274ede0d51763825

Observation b7d85cb7-4d73-4e51-bdfd-5863dc4c2aa4 · inbound

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models cites this paper.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Efficient Large Language Models: A Survey

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.959612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.959612Z digest=sha256:34d4408cfd812a48be5c24aa444f47998125ab25351967322d669d972fba4349

Observation dc37a957-0548-46d9-ab98-bf9802b373e6 · inbound

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification cites this paper.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Efficient Large Language Models: A Survey

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.428921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.428921Z digest=sha256:44da5f31b8e17aa2c4b4e78792b3670175728cef8418134f267964974a88d071

Observation 3f0d3d76-8678-430e-bfc6-829538528806 · inbound

Transitive Array: An Efficient GEMM Accelerator with Result Reuse cites this paper.

Transitive Array: An Efficient GEMM Accelerator with Result Reuse Efficient Large Language Models: A Survey

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:17.326433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:15:17.326433Z digest=sha256:c20f5504b38ce94f67fda224bffef0696df76f9e146a132599b82d4395910ba0

Observation 93d7adb2-77d5-4551-b443-cd81e3589f7e · inbound

Synthesize-on-Graph: Knowledgeable Synthetic Data Generation for Continue Pre-training of Large Language Models cites this paper.

Synthesize-on-Graph: Knowledgeable Synthetic Data Generation for Continue Pre-training of Large Language Models Efficient Large Language Models: A Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:34:25.354037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:34:25.354037Z digest=sha256:0bd388e3a58b58c2b500bf1eedd6b2243ee5679c15db43aa236b80e9ba84186c

Observation f0633638-b793-4dd2-aee3-0c481a5c102d · inbound

EfficientLLM: Efficiency in Large Language Models cites this paper.

EfficientLLM: Efficiency in Large Language Models Efficient Large Language Models: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.097404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.097404Z digest=sha256:5b0a3cde2e5fe7f46fdc8ac446508cd2e1784dc552b6578ebb7e511780e014c4

Observation f94a8adf-9117-452c-8120-defff72b0e5f · inbound

DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving cites this paper.

DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving Efficient Large Language Models: A Survey

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:31:40.679447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T14:30:47.787654Z digest=sha256:87900264bf0163681bec3d4c6aafddb229412e70e3eba524a5d6b7d120584d54

Observation 3d0d45d0-175c-4d80-830d-6149eb3feeca · inbound

ALPS: Attention Localization and Pruning Strategy for Efficient Alignment of Large Language Models cites this paper.

ALPS: Attention Localization and Pruning Strategy for Efficient Alignment of Large Language Models Efficient Large Language Models: A Survey

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:32.512737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:31:32.512737Z digest=sha256:f01f91d38e83403a6af3a03a071798537cbef44fbd76730251d5f0e12cf2cd18

Observation d6b78270-85ac-48e4-8724-6d02e4619f28 · inbound

SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling cites this paper.

SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling Efficient Large Language Models: A Survey

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:51.633694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:51:51.633694Z digest=sha256:c864cdeef2f1b024b5f43e62a6b1d8830fdc0c0b69a57a44c89df56d35167cf2

Observation 0e60339d-5a4a-41cb-b92a-c1acc6c5b01e · inbound

Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs cites this paper.

Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs Efficient Large Language Models: A Survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.095186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.095186Z digest=sha256:ad74be566b048f6577c6b00b9bbe7a6ca88e4b59f06de0de715ef1bee9c8bab5

Observation adb9d9d7-c967-4860-a65b-ae619bb15522 · inbound

NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models cites this paper.

NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Efficient Large Language Models: A Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:24.096831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:24.096831Z digest=sha256:afd4a608a5f2e173ec28f5a52024a502ef0a2d605d2f0bdfb5d6fb17d08a3f34

Observation 17a2411b-9c09-491b-81e1-c191ed6b30d3 · inbound

Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search cites this paper.

Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search Efficient Large Language Models: A Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:39.570052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:10:39.570052Z digest=sha256:66fd56ce075f955b37931d5ebe9cf792bcb1188b748b8324e44ac0ada1b99f78

Observation b0b4398d-94be-4292-9c31-18049f43db74 · inbound

Semantic Scheduling for LLM Inference cites this paper.

Semantic Scheduling for LLM Inference Efficient Large Language Models: A Survey

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T01:09:21.374055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:09:21.374055Z digest=sha256:b61c1ff786967bac5167bd8a80e2a072e3731e3cbeafe1a6affdf9451610ef39

Observation 9fedacd3-5301-4453-b94f-2a3702bf1170 · inbound

EAT: QoS-Aware Edge-Collaborative AIGC Task Scheduling via Attention-Guided Diffusion Reinforcement Learning cites this paper.

EAT: QoS-Aware Edge-Collaborative AIGC Task Scheduling via Attention-Guided Diffusion Reinforcement Learning Efficient Large Language Models: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:41.990874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:41.990874Z digest=sha256:32128bd0e4ea1748523dca511b25ed911fb628da3839f93bab8473b6d48f2cf1

Observation 957296a5-1acc-4ae2-932a-265bc358080a · inbound

TASE: Token Awareness and Structured Evaluation for Multilingual Language Models cites this paper.

TASE: Token Awareness and Structured Evaluation for Multilingual Language Models Efficient Large Language Models: A Survey

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T23:21:25.003665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:21:25.003665Z digest=sha256:d3a14c9c6e29b08d1fed918d35801f153ecf43b9d83cc1f08353e82f5986ce28

Observation 944abb55-8cab-4d86-a9a0-aec39d7abd7e · inbound

Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection cites this paper.

Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection Efficient Large Language Models: A Survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T14:33:20.422919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:33:20.422919Z digest=sha256:67e852fcf1a8ff27c4a78a32d61a925b7f52e4295be5ffa6cc522a093db8d709

Observation f98ecfeb-3a36-4d30-8512-c444a8fa88ca · inbound

DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression cites this paper.

DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression Efficient Large Language Models: A Survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:01.691397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:52:01.691397Z digest=sha256:95d127201b59e9894c10d6468953e2ab4d65d9c928f0ba918836055efbfaf971

Observation 8b0e460f-5dae-4ac0-8dbf-aa6239e85eea · inbound

MoPEQ: Mixture of Mixed Precision Quantized Experts cites this paper.

MoPEQ: Mixture of Mixed Precision Quantized Experts Efficient Large Language Models: A Survey

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T16:41:32.138621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:41:32.138621Z digest=sha256:10f374ae36157b8a78d982e94f2b7ade83f63abdc321a87e542937255f260212

Observation 07f81e60-1143-420b-a45b-ddb6c846cfe2 · inbound

NeuronMLP: Efficient LLM Inference via Singular Value Decomposition Compression and Tiling on AWS Trainium cites this paper.

NeuronMLP: Efficient LLM Inference via Singular Value Decomposition Compression and Tiling on AWS Trainium Efficient Large Language Models: A Survey

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:52:21.884286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T02:51:19.275111Z digest=sha256:7a63f3e5f6df897b62cb863e0b806a192c76d2068de295d77fabd37eb7a962a6

Observation 6a2f46ae-f827-4651-bef0-7063865becd7 · inbound

Toward Efficient Agents: Memory, Tool learning, and Planning cites this paper.

Toward Efficient Agents: Memory, Tool learning, and Planning Efficient Large Language Models: A Survey

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-03T09:21:42.553598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:21:42.553598Z digest=sha256:c4e0929695ea5bd7837eb9828df7aa4e91cbec89d5e63a2d16243deed96faab6

Observation 937bf638-f0a6-42ee-acb6-fdacb1b4cc1d · inbound

On the Limits of Layer Pruning for Generative Reasoning in Large Language Models cites this paper.

On the Limits of Layer Pruning for Generative Reasoning in Large Language Models Efficient Large Language Models: A Survey

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:40:46.217656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T08:40:24.822863Z digest=sha256:cb01788a1e0012ab2cc631a979d0f650dcfeb0c2d0ddd2d7e9d620019a8bd5c3

Observation 3b288f5e-17fc-4319-afdf-832d40031302 · inbound

Triplet Feature Fusion for Equipment Anomaly Prediction : An Open-Source Methodology Using Small Foundation Models cites this paper.

Triplet Feature Fusion for Equipment Anomaly Prediction : An Open-Source Methodology Using Small Foundation Models Efficient Large Language Models: A Survey

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:56:40.989807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T21:52:27.888915Z digest=sha256:cc8d33c1f494b6a57ea00cfb8babbc2d209161c820172128582112c9f40ad488

Observation 546db393-9dfb-4383-89c5-c9a5750d74af · inbound

CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation cites this paper.

CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation Efficient Large Language Models: A Survey

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:23:22.886951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T00:20:36.975106Z digest=sha256:6e69af176cd60e3e0bbbfe4fd4605247dafb0eb98941603a13e4f3ef1b3d0c82

Observation da83b55e-56eb-4dc7-abad-db8ccbd05f43 · inbound

Unified Deployment-Aware Evaluation of Open Reasoning Language Models cites this paper.

Unified Deployment-Aware Evaluation of Open Reasoning Language Models Efficient Large Language Models: A Survey

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:01.270847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:49:52.594972Z digest=sha256:fbab92a91439f86b617b863c2989373d31c24b5a7c6d6a2c41e97caa6be67237

Observation 82689292-8553-48f8-bda0-2e92b5a88ca2 · inbound

Unified Deployment-Aware Evaluation of Open Reasoning Language Models cites this paper.

Unified Deployment-Aware Evaluation of Open Reasoning Language Models Efficient Large Language Models: A Survey

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:44:05.726912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T09:42:33.144129Z digest=sha256:cc4190e657650f46564460687f753016890d1691d2ebc1d07c4d2d2216026ce9

Observation 0b92a5b2-d53e-4cbb-90aa-981378b81a51 · inbound

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda cites this paper.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Efficient Large Language Models: A Survey

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:30.836748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:01c23141603fe978af1d3827a0fb0e18709f63a5959baa7fff16263ed4ce27ac

Observation 3494371a-1bef-4b0d-86cf-5492daf91652 · inbound

Compress Then Adapt? No, Do It Together via Task-aware Union of Subspaces cites this paper.

Compress Then Adapt? No, Do It Together via Task-aware Union of Subspaces Efficient Large Language Models: A Survey

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:45:40.935533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T18:10:41.778201Z digest=sha256:89c5c74774c0eff465ad5ee3f575f144cf5466626622a6e2dd50f457474a501d

Observation 780515ae-7217-4263-b25d-88c6ac02bdf5 · inbound

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization cites this paper.

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization Efficient Large Language Models: A Survey

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:30:44.243755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T18:23:14.935801Z digest=sha256:db220be253150b3a0cb145af80fcc99021334f296f94e1a61315a0f842e06d8a

Observation e4086864-4fc8-4b4f-9fe9-fa9cefeb88a2 · inbound

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization cites this paper.

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization Efficient Large Language Models: A Survey

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:18.166123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T02:59:00.997742Z digest=sha256:39be2dfc0d2667637bdb667c2083592b94f5701f4c3e13efdca63dbf87a20137

Observation 77eccf07-7179-4649-bec5-97c8f97a3884 · inbound

Continuous Latent Diffusion Language Model cites this paper.

Continuous Latent Diffusion Language Model Efficient Large Language Models: A Survey

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:11:10.544287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T10:04:09.646578Z digest=sha256:1a1e15fe4520f321072d505197f9ad2d089276ffa68943e66b20558fd770c2d8

Observation 1f3178b3-5611-41c9-9c58-43cdd5293b41 · inbound

OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents cites this paper.

OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents Efficient Large Language Models: A Survey

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:27.575343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T01:57:22.554881Z digest=sha256:889dc3e50a8edb57211a229b7a34b30a9b136b8a4b3c5d32bb7573a514c44ce8

Observation 433bdf12-0e1b-4f04-b5f3-93caea3d0838 · inbound

OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents cites this paper.

OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents Efficient Large Language Models: A Survey

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:45:45.467485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T22:52:54.684185Z digest=sha256:3a661a87317aa1b42c84cd12ba3911e91e2439ee9a0c034775d1453f04e4a90c

Observation b641f171-32e8-480c-af79-a60632b11f34 · inbound

OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents cites this paper.

OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents Efficient Large Language Models: A Survey

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T05:17:31.244809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:17:31.244809Z digest=sha256:785947521371e3773a9ae98ddd80b5a4e4fb9affdf74a7b0cc54174ea95fa7b9

Observation 51e8d4f7-ff39-4b93-a88a-86c1aafbf229 · inbound

When is Warmstarting Effective for Scaling Language Models? cites this paper.

When is Warmstarting Effective for Scaling Language Models? Efficient Large Language Models: A Survey

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:02:53.104922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:00:47.126736Z digest=sha256:897c1ba7eb4c6460994fae6eff7dc2167f0fa7a21c90b5ca0b9056b2adffcdce

Observation 749b13f2-b228-4628-82b8-580fce174359 · inbound

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents cites this paper.

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents Efficient Large Language Models: A Survey

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:49:53.649954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-21T08:45:56.550821Z digest=sha256:378c2252276867285a5e083d3fc1c1ee87f4932157cdc05c38fe47091c8ecab8

Observation b1b11fb2-fafe-485f-b572-1976f07f659b · inbound

Latent Action Reparameterization for Efficient Agent Inference cites this paper.

Latent Action Reparameterization for Efficient Agent Inference Efficient Large Language Models: A Survey

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:48:12.863156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T10:45:20.306945Z digest=sha256:f4ef3688791977b0be1079c6b70718cb606cc71d9bc099aaabbbdc7bed078ae3

Observation 06f91f01-caa1-4730-9b11-03040e86f958 · inbound

EmbGen: Teaching with Reassembled Corpora cites this paper.

EmbGen: Teaching with Reassembled Corpora Efficient Large Language Models: A Survey

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:38:05.506945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T06:34:39.739666Z digest=sha256:b8f4a87dd66f8b149ac1d143b80bfc3bb6428621640c29333eca428122b10c8a

Observation cd89a295-a623-4263-a8e0-fa0fd2fc7cda · inbound

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs cites this paper.

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs Efficient Large Language Models: A Survey

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.850168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:33:06.719954Z digest=sha256:41f1d3e433da9dceb8e326901c9e4165591982eabe0458aed563a1a7737971c9

Observation 5897d9cf-d54e-4d8d-82d6-ba32d7bd5766 · inbound

Evaluation of ML Resource Utilization Requires Model Life Cycle Assessment cites this paper.

Evaluation of ML Resource Utilization Requires Model Life Cycle Assessment Efficient Large Language Models: A Survey

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:14.588198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T17:27:19.467192Z digest=sha256:7ae09f4ea24843a1ad5275fe99a623a0147ee22640e1c8d75ae30cd45bf370aa

Observation bf4af32d-1f4e-4508-8baf-1ea91200213a · inbound

Dynamic Linear Attention cites this paper.

Dynamic Linear Attention Efficient Large Language Models: A Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:27:39.795371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T13:18:18.331031Z digest=sha256:e79d1d2a0f1866d77d15465c0ae1d0f68d15731486d8688d4dd8df9981357316

Observation 8b7b1267-ff4b-4dbc-a978-b1bd9d5c8332 · inbound

Recency/Frequency Adaptive KV Caching for Large Language Model Serving cites this paper.

Recency/Frequency Adaptive KV Caching for Large Language Model Serving Efficient Large Language Models: A Survey

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:29:39.052734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T13:24:44.218871Z digest=sha256:f07e41f5e492136dc0a72a432c18fade0a4b6b59bea0adfdcc8b788933f94425

Observation 9901a9a7-6210-4689-863a-a02f9fad655a · inbound

Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks cites this paper.

Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks Efficient Large Language Models: A Survey

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:49:52.913147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T05:51:01.446139Z digest=sha256:3a4aff087ea62d26d1fc423f0addd937a949fa4a1a0471d9912ca3f60c97667d

Observation 1f5c11cf-8d01-4d77-9088-2b6d03f17395 · inbound

BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding cites this paper.

BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding Efficient Large Language Models: A Survey

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:55:41.384304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T05:59:15.199717Z digest=sha256:f01b0653679a7e57c1ffb07a209ca0dcc960edf9b7883816747bc1c3cb4cd6d8

Observation 229976d7-c436-4347-81a0-fa0ab7f87b43 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Efficient Large Language Models: A Survey

Reference 117

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T01:36:44.270544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:ec2e73bf345e2233803ff547de20259c439c5baa81367e179897b1afa3359e38

Observation b70e6bcf-a08e-4646-8684-be985470a96c · inbound

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy cites this paper.

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy Efficient Large Language Models: A Survey

Reference 245

Resolution
unresolved
no resolver link, observed 2026-07-14T06:30:16.612345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:30:16.612345Z digest=sha256:91bbdc059e245e0aa5548a03cfa2ab0bfa376686869d3e3ab74b5e502805b5d2

Observation 399498e4-c7cc-4f3e-bdcb-93978b1ec766 · inbound

Token Reduction Is Not Cost Reduction cites this paper.

Token Reduction Is Not Cost Reduction Efficient Large Language Models: A Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.741063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.741063Z digest=sha256:5cc70a2b922f7211a2886114eca748e3501e42def3dff8d009995090769c16e1

Observation db817337-0139-4590-ad76-a8058ee2b9cf · inbound

Token Reduction Is Not Cost Reduction cites this paper.

Token Reduction Is Not Cost Reduction Efficient Large Language Models: A Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T04:23:31.693317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:23:31.693317Z digest=sha256:7ae09bc5e768012479aebe9e258dea5fb88ddfb3be55238a83ae6905a75cf90d

Observation 7f446d23-3496-4481-ab49-d225a2486905 · inbound

CHS-SQL: A Text-to-SQL approach based on Confidence-Guided Heuristic Search Schema Linking process cites this paper.

CHS-SQL: A Text-to-SQL approach based on Confidence-Guided Heuristic Search Schema Linking process Efficient Large Language Models: A Survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T11:00:12.321867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:00:12.321867Z digest=sha256:b8741eb0ed1ee55d66f6f1b1b06b0f3fc40c461026b44d99c1313104583f39b6

Observation ec33ef9f-be68-476f-8a3e-c89ac65c5000 · inbound

LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs cites this paper.

LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs Efficient Large Language Models: A Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:55:56.385913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:55:56.385913Z digest=sha256:9e4e0fe49ee81c378ac788766e191ccbdac6383b8c71b3a67bd43f21fe2077ec