Pith. sign in

Paper Citation Record · LEDGER

Efficient Large Language Models with Zero-Shot Adjustable Acceleration

As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2509.01190.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01190 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:52:44.664299Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bfa35c9f-fdef-4003-9f08-d30a3be4ddf8 · outbound

This paper cites LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.609841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.609841Z digest=sha256:4106b9b60b3d3b43c4016fd9463859e9a9e18267500ad0b88796d05afb19450f

Observation 9617c512-fa86-40b4-ac56-e644275340ea · outbound

This paper cites SmartBERT: A Promotion of Dynamic Early Exiting Mechanism for Accelerating BERT Inference.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration SmartBERT: A Promotion of Dynamic Early Exiting Mechanism for Accelerating BERT Inference

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T12:52:44.760293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:52:44.629911Z digest=sha256:1cd010c4a13f8c261375c68405afb3487b864b081222d313cc1d9f99fe213e0b

Observation fe5ab983-814a-450f-90b0-d9eca33518af · outbound

This paper cites Heejun Lee, Minki Kang, Youngwan Lee, and Sung Ju Hwang.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration Heejun Lee, Minki Kang, Youngwan Lee, and Sung Ju Hwang

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:52:45.011281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:52:44.636353Z digest=sha256:737ce3010a521d8d31f5610c450a9c26682622cb8828d1992d10f6b638e6708c

Observation d5bd6782-a633-4b16-9a0c-2db2ab7e5628 · outbound

This paper cites RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language Models.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:52:44.721825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:52:44.641875Z digest=sha256:496cdad490aec19859c7206355642ed3ef1fad09492297df5c8ecc3c4c911ed2

Observation 3e917123-a917-467e-bf33-8a4af29a9b71 · outbound

This paper cites InProceedings of the 16th conference of the European chapter of the association for computational linguistics: Main V olume, pages 91–104.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration InProceedings of the 16th conference of the European chapter of the association for computational linguistics: Main V olume, pages 91–104

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:52:44.971096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:52:44.654033Z digest=sha256:2ff532ce33c2c669b9796568a7cb8f127297c8630915fdff1db7b43366835a13

Observation cf8f6fb1-49bb-4c5b-b2fd-b5ac5bbf8d0d · outbound

This paper cites In 2023 60th ACM/IEEE Design Automation Confer- ence (DAC), pages 1–6.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration In 2023 60th ACM/IEEE Design Automation Confer- ence (DAC), pages 1–6

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:52:44.938126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:52:44.658328Z digest=sha256:055308aaeb89aedec570c57d8ce105739fb07642c35f0640464f253d50413397

Observation 107c0efa-055e-4229-8830-5e8482b3c100 · outbound

This paper cites Blue squares represent preserved tokens, white squares represent pruned tokens, and the red line indicates the overall preservation trend per layer.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration Blue squares represent preserved tokens, white squares represent pruned tokens, and the red line indicates the overall preservation trend per layer

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:52:44.917986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:52:44.664299Z digest=sha256:d1255b0a5509b0e25d496c8864052202dc43697d9f9413ec8d4b06a73c2273b7

Observation c16095b2-c899-494f-9a47-984c45f04b98 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.604653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.604653Z digest=sha256:b6546a774533aab051b29a5a2669073376c6d4e3cf21053ebf05f6a9cd6c6e03

Observation aecc832f-0f19-45a3-b2b6-e1f0bdcfac3d · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration Measuring Massive Multitask Language Understanding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.623223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.623223Z digest=sha256:2cfce3303858b2b23ca46da41e87f1e8e77e1f0ad28bc17670d8737b926ce582

Observation f0b1ab8e-093c-472a-a808-1026bf03a6c1 · outbound

This paper cites InICASSP 2021-2021 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7713–7717.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration InICASSP 2021-2021 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7713–7717

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:52:44.992848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:52:44.649209Z digest=sha256:bdf842af84d8f043baf61f63ce46af85c44f627f26856da5c0975b449e89e2c6

Observation 379ad255-dff5-4788-8671-cd5f7f99c0d7 · outbound

This paper cites Token Merging: Your ViT But Faster.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration Token Merging: Your ViT But Faster

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.591356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.591356Z digest=sha256:caf8eb191b171aa48c338aed42a7a96f99cba8c37a9d6110677e9bf9667a86cc

Observation 1eed93bd-164e-4734-afbc-880163b626f3 · outbound

This paper cites LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.617762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.617762Z digest=sha256:3ce877d9b691dee26add960d4f5d4c4bb10d62df8171e1654fbc2ab0645440d5

Observation 8bfc1734-8aa2-4359-a406-3592205cbe19 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.599095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.599095Z digest=sha256:ab0abcc0590db4e4004e301c4ada8104de93a2df71ed19b0272d8aa55a973964

Pith citing papers

No inbound Pith citation observations are available.