Pith. sign in

Paper Citation Record · LEDGER

LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2407.14057.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.14057 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:49:21.976799Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T00:56:40.961590Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f1e8f27e-4c36-4d5b-a8dd-e8e522e1cc1b · inbound

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding cites this paper.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:21.976799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:21.976799Z digest=sha256:f102f629df1d32fcf1f2aeec2e065d0aace9b3784f13f0d7c7bc208460b2c940

Observation e8b7dab3-7667-4bb2-91bf-6d44d6c44b0b · inbound

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models cites this paper.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.633017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.633017Z digest=sha256:bcf1348e0ccdd1c049ca07f692b7817014201b4226a0bf6bcf0447885119e8db

Observation b58ca8af-8e8a-47d0-b4fc-e6a66fc6ec52 · inbound

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning cites this paper.

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:55.187485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:55.187485Z digest=sha256:49dfebe3a274dc92103f3204e859815c46ca99728b26f7e893c11e76f6ee9555

Observation d6cf978e-0eb9-44c4-9592-8fb91282da87 · inbound

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity cites this paper.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.998574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.998574Z digest=sha256:bbee13fafa08835f7c495247ebee8fd13b91b51ad663d395025bb9fcf82d9299

Observation 88ca61d6-2467-463d-9ec8-119d1d6d4566 · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.056197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.056197Z digest=sha256:ea4b667489bd2a925d8bbffcb9722856bfcda1cfdbbf0b9c54dfa52207fe9b49

Observation 16030a24-627d-42f6-9656-a4b7149b6626 · inbound

DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs cites this paper.

DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:57:17.905006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:57:17.905006Z digest=sha256:6390ec4d778f0c97495d0c917011f70515655f56e23772bd0a7e8b32460b71d4

Observation a127aadd-1e90-49a8-ac08-bc2d54c11073 · inbound

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models cites this paper.

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:24.993164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:24.993164Z digest=sha256:9cd2869fc50903a67f5bf208528b30462b689715b63f0a9ed01ca7b66ce1caa7

Observation 906151df-18d8-44de-a0b1-6f5b768d23fe · inbound

AdaFV: Rethinking of Visual-Language alignment for VLM acceleration cites this paper.

AdaFV: Rethinking of Visual-Language alignment for VLM acceleration LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:00:12.593109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:00:12.593109Z digest=sha256:972decd4eea6dbaeb7313537deb6fb5cbc63d1c6a047780551e1f711d825ed23

Observation 6ddf8048-7b81-4ad7-b994-1d2c9de8826e · inbound

Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference cites this paper.

Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T16:40:36.625437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:40:36.625437Z digest=sha256:7a1d9e82ea87274f5fdc1ab31e3c38f5b227ba8695eef87956585cec766cf51f

Observation 26e91560-350c-408d-9815-e258f90218e4 · inbound

SwiftPrune: Hessian-Free Weight Pruning for Large Language Models cites this paper.

SwiftPrune: Hessian-Free Weight Pruning for Large Language Models LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T15:21:23.439544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:21:23.439544Z digest=sha256:dda86e170ec60b44229c9bd99de763fbd84f3c9bd7ad0a3786f75a3f73a30731

Observation 62bc1830-6ac5-4227-ab28-144eedf6fe5e · inbound

TransMLA: Multi-Head Latent Attention Is All You Need cites this paper.

TransMLA: Multi-Head Latent Attention Is All You Need LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T11:48:01.188051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:48:01.188051Z digest=sha256:d45ab9a489bd1ae999330c2d3cbf0db5ce233ac25930a5bf11896fa40866322f

Observation f1e4bc0d-81d0-4354-bc7d-ea55a0ebc2c7 · inbound

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU cites this paper.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.822448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.822448Z digest=sha256:0a3381cc5beea08779b7ee3fdcffe6e13f3906b96e2cebd79e84af22a2411a28

Observation 7a42e937-f4c5-4818-95b8-56574eddf90d · inbound

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache cites this paper.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.925945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.925945Z digest=sha256:94493abce00575822d46b54e87a63bc7e998a14744c33afe04638fd4bc7e0c94

Observation 241ba12b-38b3-44e0-975b-b543ce66e8a2 · inbound

Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration cites this paper.

Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:43:24.234222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:43:24.234222Z digest=sha256:2dd0dbb267efc5b0c17a004ba6ac1705a74842f4631d69cc44675096ef18d50d

Observation d6750712-6789-43a0-8b92-61b36e8e0c86 · inbound

Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning cites this paper.

Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:55.004748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:32:55.004748Z digest=sha256:76b06eb93fe3cf40c541d0f361fb178c7af69423b06af8006ea4ac450d7e2cd4

Observation d7ec6af8-00c3-42e1-aba7-a1a71f7b5545 · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:08.556414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:08.556414Z digest=sha256:0c9d73c939743f7af976cb35c9fd38a321cc21c1447694ca1442a363f7b6db26

Observation bfa35c9f-fdef-4003-9f08-d30a3be4ddf8 · inbound

Efficient Large Language Models with Zero-Shot Adjustable Acceleration cites this paper.

Efficient Large Language Models with Zero-Shot Adjustable Acceleration LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:44.609841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:44.609841Z digest=sha256:41f79c190cb30b1bc2450588a5cd956d8747f00cf266784383526ba2897b30a7

Observation 095b1f06-f40b-4e43-a394-3b7c995d4bf2 · inbound

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems cites this paper.

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:15:06.802670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T12:14:09.509302Z digest=sha256:0b80aa9005bbafd4e19c72ed2bc6b05ade6c2e39ebff0890006620872fb2843a

Observation b2c48ef6-0535-4acd-81ab-175191fe5533 · inbound

Three non-Hermitian random matrix universality classes of complex edge statistics: Spacing ratios and distributions cites this paper.

Three non-Hermitian random matrix universality classes of complex edge statistics: Spacing ratios and distributions LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T16:19:49.419805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:19:49.419805Z digest=sha256:a8ba843d36cb34a6399cae3adb13b1a7fa0365633f32d553cb5fff39d2887e74

Observation b450e870-4055-41ef-b915-f0328dd08015 · inbound

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention cites this paper.

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:43:00.779043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T21:40:28.854082Z digest=sha256:b6f3b689b2011e5186bb4c12946d2fd2bfc3e26dc95a44a100f69e4432f8b057

Observation 61ba891d-7883-404b-b8ab-d92c297cd844 · inbound

UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification cites this paper.

UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:56:08.366548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T10:43:01.724760Z digest=sha256:d33894c160afebe13d81f7ae770b7002577312ef5af86b694db1eb76a1c73feb

Observation 432516dd-b0d1-43bd-a336-0787e7da1fb6 · inbound

Long Context Pre-Training with Lighthouse Attention cites this paper.

Long Context Pre-Training with Lighthouse Attention LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:11:09.463896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T10:10:38.610613Z digest=sha256:be3cf468067cc6463daf1864c8760343a1ec02047d35072dd1780fbfe96f0fcb

Observation 13087462-aed8-413f-922b-2539c831d469 · inbound

ProactiveLLM: Learning Active Interaction for Streaming Large Language Models cites this paper.

ProactiveLLM: Learning Active Interaction for Streaming Large Language Models LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.874551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T19:07:48.425181Z digest=sha256:8b82c7f3b8c424e78a5b646dd50d174f5d4e89cd6da24e1db5fbee707857b8e0

Observation 037ea3a9-2cf1-49e0-a4c6-45d98885b7a1 · inbound

Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization cites this paper.

Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:56:40.962769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T00:55:52.215700Z digest=sha256:82338c70eb0e722134eb76664f30db24b90f8847caca12f00643075d59bc2c88

Observation 2518dda7-f2aa-429c-ac6a-066eadd27f77 · inbound

Structured Thoughts For Improved Reasoning And Context Pruning cites this paper.

Structured Thoughts For Improved Reasoning And Context Pruning LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T12:08:05.502310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T12:08:05.502310Z digest=sha256:8e2431b544f05dd2488b9a073f10752abed1c6dda22146c60f59699a85a96a96

Observation 421cfbb6-0894-4c99-9f39-701ab756934d · inbound

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling cites this paper.

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T09:51:03.373817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:51:03.373817Z digest=sha256:00288d6110bff7ebf04d896108495c1a9ebc1664c57d08695111f8b9e31bad0d

Observation 8c003081-64d6-4499-944b-ac596e9bfbf5 · inbound

Hierarchical Domain Generalization cites this paper.

Hierarchical Domain Generalization LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T20:54:08.630899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T20:54:08.630899Z digest=sha256:3c89ccf053bc289324756bf54712e44ba8e06fd27382bc6e2b65183eaac69501

Observation ec0f7d09-cb78-4a40-a4ec-8285c62a8ff6 · inbound

PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention cites this paper.

PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T11:04:04.538729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T11:04:04.538729Z digest=sha256:bb829a3d50fcabad0793b1bb8e7d14d2103e1f96512626b64470103c39d8d452

Observation 7abf6c4e-4cbc-41bc-97a8-00d91866cd5b · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.394782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.394782Z digest=sha256:258f79060d5550f461a38c4198c639c22c2d6c1eb55edf249e4e95fef99e88ec