Pith. sign in

Paper Citation Record · LEDGER

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking

As of 20 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 3 inbound Pith citation observations for arXiv:2412.01380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01380 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:30:36.140273Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:15:49.289390Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T13:21:35.626449Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99550ca1-e45d-4b67-b230-96d096becd2c · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.868916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.868916Z digest=sha256:0f2fac58f6b65501df46f8d186b3978529c484d3b171dbd25ce1c108d9b2eae3

Observation 9407622e-fc95-4455-b73a-5bfdd1865fa4 · outbound

This paper cites ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.874693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.874693Z digest=sha256:2515d8bb0fd85f8e598533dab9c6bade711f1735ab6bcab41bad0d46e3610ae7

Observation e1733bd0-7572-426d-870d-4a6a60f7c7ac · outbound

This paper cites LLM in a flash: Efficient Large Language Model Inference with Limited Memory.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.880444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.880444Z digest=sha256:c4ae3fbf734afaf5b047110cb165ccebb73b2f461509765dbe5c230bf369e312

Observation 54f6ffc3-c380-4cef-9e91-386af7bbcd0f · outbound

This paper cites Qwen Technical Report.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.885851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.885851Z digest=sha256:27b8c3e23d9762b00fcfab94a6cef825f1ccd8f469985d1fa6b808cf27b9fd09

Observation cc1078dc-490d-4158-8169-e5a88d67312c · outbound

This paper cites an unresolved cited work.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:30:37.255188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.890455Z digest=sha256:286ae964b943c594c34b1990c9d76d86fd38b84758504b43b637579abf78e823

Observation 845a17a7-00dc-4003-81bc-f2a6868c3fdb · outbound

This paper cites Smartphones beat dram drum to meet performance demand, 2021.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Smartphones beat dram drum to meet performance demand, 2021

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.239712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.895668Z digest=sha256:090b3a457d27f433a462f62186caec93ae0648cecfcf5fa62714e57cecc08329

Observation c977e7a0-181c-496e-9c1d-4172fe11b670 · outbound

This paper cites How much ram should a phone have, 2023.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking How much ram should a phone have, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.220848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.901181Z digest=sha256:a0dc0a0f6f567d4d34dbe77b28cd3212735bd5996f7a4f2c86a7835eb21b9460

Observation 6f33b5aa-e1b5-4aa1-91f2-422074d7cdd6 · outbound

This paper cites N., Fan, A., Auli, M., and Grangier, D.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking N., Fan, A., Auli, M., and Grangier, D

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.905750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.905750Z digest=sha256:fb404510aec32bbb46428aec3326c60eff62d42b3e323131ebe9481ee94b2612

Observation 54d12ab3-7d07-47d2-832b-b0e20c642ee6 · outbound

This paper cites The Llama 3 Herd of Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.910387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.910387Z digest=sha256:fa1081af79b447185372af07ea37c7621086ea1df588f09d42e4256f19fa1999

Observation aa89924d-8db7-4aa8-8592-bfca4782e7e7 · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Sigmoid-weighted linear units for neural network function approximation in reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.915251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.915251Z digest=sha256:6bfc3a56f1ef305e8352fe4a91ac60714ff6c53faf8f6b06ec3b9730e5111a1f

Observation 53c840b7-2ac2-4992-bc00-60d1e7fde47f · outbound

This paper cites Code artifact for efficient llm inference using dynamic input pruning and cache-aware masking, March 2025.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Code artifact for efficient llm inference using dynamic input pruning and cache-aware masking, March 2025

Reference 11

Resolution
verified exact
doi, observed 2026-08-12T04:30:36.183513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.919920Z digest=sha256:7939903db094a3920fe91b57d6664bf84aab31dda88998ca694198c258547216

Observation c96977b5-55f0-4e86-adc9-c0a463f85676 · outbound

This paper cites and Alistarh, D.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking and Alistarh, D

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.924610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.924610Z digest=sha256:b19d4148584daf2042bfbd779d8a65fdce82e37f4abfa6a39f77bc6a073c6eca

Observation c5a7f179-82da-4fc0-bfa6-429a915aa293 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.929393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.929393Z digest=sha256:0c099a4a6c5bb1353b27f53e60889cbc92bd6689c2b9f3df0b3e1458c6d808e1

Observation 0c69a890-7a6f-436b-a286-0c27c9a20853 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking A framework for few-shot language model evaluation, 07 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.934418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.934418Z digest=sha256:d0618b74afce9789106cacf3b3fb43e2dfd291a4624438c07d8640b4f6d49182

Observation 9a84b3fb-b85c-4189-829e-062bea9a86f5 · outbound

This paper cites W., and Keutzer, K.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking W., and Keutzer, K

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.938996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.938996Z digest=sha256:9df1265d1ed06ae765fd816011924da780fbd9eafc1affde4871cbadb060241a

Observation a52cae37-1f11-4309-9d1f-2f0678d8cf28 · outbound

This paper cites and Lorenz, J.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking and Lorenz, J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.161924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.943437Z digest=sha256:1170b2603eb83af6873955087c1f15ba10e2de7eeaab7b335e96137bf88cbc02

Observation de89ce20-02e1-44f0-b1f1-3b77bbc824a1 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Measuring Massive Multitask Language Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.948004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.948004Z digest=sha256:0b0cd634d363a13bd4c220fd22f6f43c19854c7dea456463b2d00d158becef78

Observation c159906b-805b-451f-8773-775215900693 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.952779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.952779Z digest=sha256:6b0383132b4de780ab2fbc08fbeece4e1c2d67624448fd66f011849765a3f12c

Observation 980c6f8d-38d7-4218-89d8-32e703c6c764 · outbound

This paper cites BiLLM: Pushing the Limit of Post-Training Quantization for LLMs.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking BiLLM: Pushing the Limit of Post-Training Quantization for LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.957853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.957853Z digest=sha256:0416e57285857033a4c6c83ca69c99ab5ae68652cda7da71e3d137e63b68170b

Observation 96bcb492-48e7-4f5d-9a59-c9ddd76ae092 · outbound

This paper cites and Lin, C.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking and Lin, C

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.146110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.962685Z digest=sha256:bd33f6e57fad46ef13ffcca6a245bb370c66e48123fb46029aa45f723025ac16

Observation 7b972b95-ad44-4721-912b-66af0df3758f · outbound

This paper cites Challenges and trends of sram-based computing-in-memory for ai edge devices.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Challenges and trends of sram-based computing-in-memory for ai edge devices

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.130128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.968421Z digest=sha256:1fdbb997232d339fa4a76114364d8c5f754f103f95ff3174d3fa513b5ca69d9a

Observation 1ffd6fdc-b3bc-4870-bc4f-035d3326d324 · outbound

This paper cites Mistral 7B.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Mistral 7B

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.972943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.972943Z digest=sha256:3aa4ee61b44f3192b125680a42f0ea00d6e23fa860f4bd5f852207e95c514831

Observation f1b47f41-f004-43be-98c0-9698e15ca0fc · outbound

This paper cites Pruning vs quantization: Which is better? Advances in neural information processing systems, 36: 0 62414--62427, 2023.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Pruning vs quantization: Which is better? Advances in neural information processing systems, 36: 0 62414--62427, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.113815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.977465Z digest=sha256:528d6e9624db6635ca89db81ddb7783f2ec8362cf3da355af3053f8bacc041fc

Observation f0f7afbc-d75c-482f-8ae6-ecfe48b82a42 · outbound

This paper cites Pruning vs quantization: which is better? Advances in neural information processing systems, 36, 2024.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Pruning vs quantization: which is better? Advances in neural information processing systems, 36, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.096028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.982024Z digest=sha256:5f11fb8860343e11402eb57a79d80aa90807209f9b161e6c9c718b7dd00fa819

Observation 347859e4-90cd-417d-9aaf-18d7d12244c1 · outbound

This paper cites H., Gonzalez, J., Zhang, H., and Stoica, I.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking H., Gonzalez, J., Zhang, H., and Stoica, I

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.987620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.987620Z digest=sha256:18c38cfd6fcef3f61c9a79657c8f6ea9236db5953de911b69b043c457a7230e3

Observation 6a1c6410-ba49-4e31-994f-f5ca55b5f7af · outbound

This paper cites Apple nvme vs ufs 4.0 storage, 2023.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Apple nvme vs ufs 4.0 storage, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.068082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:35.992271Z digest=sha256:c8b05dfd408e1e20c8016ceab6d38888951710666fd500e3dc688d0944393b48

Observation 5d5c0d83-edb6-4a88-b3f4-a6c401ad461e · outbound

This paper cites CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:35.996733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:35.996733Z digest=sha256:19b3a38962cd2d3e051e73082025ed440b5179f7f1da2d2d10826524a0fd8c04

Observation a047825c-5268-4438-aacf-4ff4eee851a3 · outbound

This paper cites an unresolved cited work.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:30:37.050115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.001626Z digest=sha256:555426fb4832cbb53a3bace5842755fc18ca74bb56f145323782cf154189a3d4

Observation ea440abe-924e-4ad6-9b97-e15a61138d6f · outbound

This paper cites Deja vu: Contextual sparsity for efficient llms at inference time.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Deja vu: Contextual sparsity for efficient llms at inference time

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.006409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.006409Z digest=sha256:a203ad69bd968fbd48fd537216bba9d932f0d94e861a771ce710276e013b149c

Observation 40c17f66-d939-4db3-a385-2538d8109a1c · outbound

This paper cites and Vassilvitskii, S.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking and Vassilvitskii, S

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.010972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.010972Z digest=sha256:8ebf5a9d0e158b98066279a56f9aa1824681ed83185207ed3274f53de8068b3a

Observation 661d8d82-7e10-4393-8148-858695803850 · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Llm-pruner: On the structural pruning of large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.015725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.015725Z digest=sha256:866bd4e1354acc2430f0ce586d9f72762486fcd4592885dd6aedc3c4c7bffdc1

Observation 33594af2-a8c9-4ee6-8f9b-88a2eb90954b · outbound

This paper cites ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.020982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.020982Z digest=sha256:45169dc0c539bbc1b21ad4407f21dfa5bd209f8573a815afb5b56745c77618c4

Observation 54c22a46-6e6f-4636-979a-6258aa59bcf4 · outbound

This paper cites A White Paper on Neural Network Quantization.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking A White Paper on Neural Network Quantization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.026374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.026374Z digest=sha256:c9cb077084da16b1dbe3c5ba0ae775fdc7398d95ec7c4ef3fa60dc5a07cfb030

Observation 77168a8e-e847-4c47-bc55-aa7753c693a5 · outbound

This paper cites A Comprehensive Overview of Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking A Comprehensive Overview of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.031099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.031099Z digest=sha256:986de75f7c0e4f16838da1f0f70706a4b23e727801e535b26f1fe5b4fd918f4b

Observation e87065d3-03dc-4e94-aa3c-bf575b49cace · outbound

This paper cites Llama 2: Early adopters’ utilization of meta’s new open-source pretrained model.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Llama 2: Early adopters’ utilization of meta’s new open-source pretrained model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:37.001568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.036146Z digest=sha256:0b87f3a9156a61b58ec364facfd6a906a4a6d4519e455962714073f7a558c8e3

Observation 0808ae4c-4854-4cf7-815b-413bff1b881e · outbound

This paper cites an unresolved cited work.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:30:36.986475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.040734Z digest=sha256:4bfb26ff35fda535ef2fd755931d075ab77f1cc97f8492016085bf47d59c18be

Observation b564d181-3870-4d71-91a6-bba295237c1c · outbound

This paper cites Effective mimicry of belady’s min policy.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Effective mimicry of belady’s min policy

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:36.970840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.045822Z digest=sha256:10270a5e6cb4368b6418721cd9ae3bd0c4581aacdf2ecddcd072197c6ab78ad9

Observation a1f17811-df1d-40ef-b74a-c28de383297f · outbound

This paper cites R., Hestness, J., and Dey, N.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking R., Hestness, J., and Dey, N

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:36.955021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.050610Z digest=sha256:4d0a5655ac77d11df7a85a974fd71f3fe742a041dd9324957b6873827f3f8147

Observation 9653c257-04d8-4d09-be4a-8622a7579b98 · outbound

This paper cites ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.055200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.055200Z digest=sha256:f8691cca4c46bf20b33e6f64fe9ad0cc7157f64d3a323ff46dc637539d2bc7a2

Observation 0727e386-f3a9-497b-8cf2-561e39e4bada · outbound

This paper cites PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.060140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.060140Z digest=sha256:56a531e3425f6eb9409cb9e95bbd6c65d37e88b88a68319b5daee47dae6cdea0

Observation ba622ce4-2691-4519-87dd-712870577f6a · outbound

This paper cites Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.065313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.065313Z digest=sha256:370c3b482ccf7a51debb5f274f771b597f3fc39c13aab252e51cf62cdb2f4bf2

Observation bb59a358-0219-40a4-8d81-8de302534ebf · outbound

This paper cites Sparse language models with relu activations.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Sparse language models with relu activations

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:36.807249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.070513Z digest=sha256:ae19858fdf7cb5d23aa421b77f15f59f0063ddfd152e1c320a2a07bc2a7b2357

Observation a0031b78-b421-4dd9-8280-2c322a9c32ae · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking A Simple and Effective Pruning Approach for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.075170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.075170Z digest=sha256:7a5ce2812c60871c50bd34373e83c3e63d25ae2d95defa633f3cab93bf9f9af6

Observation 3088ac87-2ccb-4896-9e5c-2459bfbdc331 · outbound

This paper cites GPTVQ: The Blessing of Dimensionality for LLM Quantization.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking GPTVQ: The Blessing of Dimensionality for LLM Quantization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.080091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.080091Z digest=sha256:7254099275dfac093e46c5cfb840cb4a7cd7afc9ae3a242bc3c59c742a07dafb

Observation 6655a9d5-9afc-413a-8fd7-e7ce627e7381 · outbound

This paper cites The LLM Surgeon.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking The LLM Surgeon

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.094555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.094555Z digest=sha256:7fb793c53be743fff6e05a03aebb2d84d0d5fc106fd8a0eea6d2393e51e4f63f

Observation 170d6ec3-2d9b-4302-848e-39d0cc1729f8 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:36.790571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.099571Z digest=sha256:5d15f50566decd435b799785a6738305c780be7042edc95170661b388d1ab560

Observation 9bc7e5cb-0308-46eb-8502-a7c4732ba1c3 · outbound

This paper cites Apple silicon --- Wikipedia , the free encyclopedia, 2024 a.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Apple silicon --- Wikipedia , the free encyclopedia, 2024 a

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:36.773992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.104470Z digest=sha256:aa83d85c3153c7753a32a5754809d12809e8c1025cefa143bc0320027c615d43

Observation 33a9ae10-e9ca-4ed8-98e0-17fd196aa542 · outbound

This paper cites List of qualcomm snapdragon systems on chips --- Wikipedia , the free encyclopedia, 2024 b.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking List of qualcomm snapdragon systems on chips --- Wikipedia , the free encyclopedia, 2024 b

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:36.754752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.109317Z digest=sha256:478400000eb4b693e78df92b7220ae55048a75c30de1d68fdf56dfe04883dc9c

Observation 3318d1f0-5035-4b36-9cf9-243395fd3c4e · outbound

This paper cites Flash memory --- Wikipedia , the free encyclopedia, 2024.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking Flash memory --- Wikipedia , the free encyclopedia, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:30:36.738500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T04:30:36.114390Z digest=sha256:ecca67e57484d44500287bd4e28d2eb704eab9da52d7dc5b24c8ac082dea1e5c

Observation 8c7d2330-15c8-4845-b7b3-f9906a84d1b9 · outbound

This paper cites PowerInfer-2: Fast Large Language Model Inference on a Smartphone.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.119613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.119613Z digest=sha256:f4b79296bfb044578a88242fdc2116be86f16b9381011e6dd32781502274fa11

Observation 08c025d0-cc7a-4893-b154-3e362a848f6d · outbound

This paper cites OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.124872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.124872Z digest=sha256:8b667bbfcba9de7d802b85b7a828c288ec1223af5b88b5b1401b3a2187b87195

Observation 31f10859-551d-4b20-81e3-9e7caedddce9 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking OPT: Open Pre-trained Transformer Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.130490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.130490Z digest=sha256:7dcae7248d4fefcec8631cd6427670f3e06ab9f58eebc92aa798f55031c346f5

Observation 667f0a38-94ac-477e-93b1-340159c08a00 · outbound

This paper cites A Survey of Large Language Models.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking A Survey of Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.135498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.135498Z digest=sha256:58fc6340540b89f19b608dcd59d4b6c627aba4b9adcf4a289b90b6300745c887

Observation e3d4b92d-51ca-4b6d-b01c-ea9de18e957e · outbound

This paper cites write newline.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.140273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.140273Z digest=sha256:873cfd133378d785de3fa34fe34df923b19971a9e9d930db4fc321e2dcd02d55

Pith citing papers

Observation 57e8b28b-7346-4cbf-9aca-f48e6a169591 · inbound

RAP: Runtime Adaptive Pruning for LLM Inference cites this paper.

RAP: Runtime Adaptive Pruning for LLM Inference Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:21:35.629058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T13:20:41.739571Z digest=sha256:a811995dda64159fd5a92f079fd33256d8c5d99989033210d553ddd67c11dc74

Observation 244a3fa0-3f99-483d-b108-e6139357ad16 · inbound

Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures cites this paper.

Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:15:49.289390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:15:49.289390Z digest=sha256:8f723a9df4f671631b7f98f9a9b700cc89da2f4dfbab4a326bc02567ef5c91a5

Observation b1490825-f9dd-4591-a84c-a20a207d8aab · inbound

AutoNeural: Co-Designing Vision-Language Models for NPU Inference cites this paper.

AutoNeural: Co-Designing Vision-Language Models for NPU Inference Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T18:56:59.496372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:56:59.496372Z digest=sha256:29ac52040ffa65846acb3919737ef1bd674abc63ec34d34b7ede88f071f5eaff