Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2407.14057.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:49:21.976799Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T00:56:40.961590Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation f1e8f27e-4c36-4d5b-a8dd-e8e522e1cc1b · inbound
Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8b7dab3-7667-4bb2-91bf-6d44d6c44b0b · inbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b58ca8af-8e8a-47d0-b4fc-e6a66fc6ec52 · inbound
CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6cf978e-0eb9-44c4-9592-8fb91282da87 · inbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88ca61d6-2467-463d-9ec8-119d1d6d4566 · inbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16030a24-627d-42f6-9656-a4b7149b6626 · inbound
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a127aadd-1e90-49a8-ac08-bc2d54c11073 · inbound
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 906151df-18d8-44de-a0b1-6f5b768d23fe · inbound
AdaFV: Rethinking of Visual-Language alignment for VLM acceleration LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ddf8048-7b81-4ad7-b994-1d2c9de8826e · inbound
Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26e91560-350c-408d-9815-e258f90218e4 · inbound
SwiftPrune: Hessian-Free Weight Pruning for Large Language Models LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62bc1830-6ac5-4227-ab28-144eedf6fe5e · inbound
TransMLA: Multi-Head Latent Attention Is All You Need LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1e4bc0d-81d0-4354-bc7d-ea55a0ebc2c7 · inbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a42e937-f4c5-4818-95b8-56574eddf90d · inbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 241ba12b-38b3-44e0-975b-b543ce66e8a2 · inbound
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6750712-6789-43a0-8b92-61b36e8e0c86 · inbound
Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7ec6af8-00c3-42e1-aba7-a1a71f7b5545 · inbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfa35c9f-fdef-4003-9f08-d30a3be4ddf8 · inbound
Efficient Large Language Models with Zero-Shot Adjustable Acceleration LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 095b1f06-f40b-4e43-a394-3b7c995d4bf2 · inbound
PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b2c48ef6-0535-4acd-81ab-175191fe5533 · inbound
Three non-Hermitian random matrix universality classes of complex edge statistics: Spacing ratios and distributions LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b450e870-4055-41ef-b915-f0328dd08015 · inbound
HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 61ba891d-7883-404b-b8ab-d92c297cd844 · inbound
UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 432516dd-b0d1-43bd-a336-0787e7da1fb6 · inbound
Long Context Pre-Training with Lighthouse Attention LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 13087462-aed8-413f-922b-2539c831d469 · inbound
ProactiveLLM: Learning Active Interaction for Streaming Large Language Models LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 037ea3a9-2cf1-49e0-a4c6-45d98885b7a1 · inbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2518dda7-f2aa-429c-ac6a-066eadd27f77 · inbound
Structured Thoughts For Improved Reasoning And Context Pruning LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 421cfbb6-0894-4c99-9f39-701ab756934d · inbound
SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c003081-64d6-4499-944b-ac596e9bfbf5 · inbound
Hierarchical Domain Generalization LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec0f7d09-cb78-4a40-a4ec-8285c62a8ff6 · inbound
PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7abf6c4e-4cbc-41bc-97a8-00d91866cd5b · inbound
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.