Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T08:00:58.889699Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2607.28418.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T08:00:58.889699Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 14a5e14e-eb41-47ba-ba23-52a3b4de683e · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Check- points
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28e3d8ed-21e3-46ad-a6e6-ef790d2c7e65 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Aayush Gautam, Mukul Gagrani, Junyoung Park, Mingu Lee, Chiris Lott, and Narasimha Reddy
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 381875ad-a1bb-49c5-90f3-3719cf8d460a · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10e4d762-f8a0-43a9-91d1-7b22c387d452 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b69cac-ce80-43e4-a590-29f0286678a3 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning What Matters in Transformers? Not All Attention is Needed
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1549c1fe-7474-49c3-982e-86d91dba9b22 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c213ac60-5bda-4ec8-91fd-8196332a3d9b · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Deterministic Differentiable Structured Pruning for Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 495e1156-fba4-46a6-b714-4b891f85dcdc · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7989027-7692-42dd-8eed-787777d7179d · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d3b58ca-6464-4a25-be78-e48ffc7d3e96 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning ISBN 979-8-89176-256-5
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5606eacd-749e-4cea-95ae-268e8308412e · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8917b35a-ef5c-47fd-b702-a68d36f06791 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38989719-78ca-4c0e-b31c-7faf68d8168f · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Kimi K2.5: Visual Agentic Intelligence
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d115edfb-a0ad-44d3-b67d-fd5229029cd3 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a565838b-d42b-4e62-9de7-fa3f45a0f6da · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning ISBN 978-1-4503-6719-6
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73e7ee6a-a361-482d-84aa-7685fb4d7dcb · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning From data to model: A survey of the compression lifecycle in mllms
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f8c2fc2-b59f-4750-9b50-d251b67ba0b6 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc606566-4ed4-4758-9a86-1855bb5c6cdb · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning SGLang: Efficient Execution of Structured Language Model Programs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff557a35-e2cf-4d95-b30c-1011573e922a · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning BlockPruner: Fine- grained Pruning for Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb712a9d-6a89-4271-91eb-9e7755f88ad2 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning mk,mnk->mn
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebed3562-4213-4a99-9420-a259a3cc4ef5 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning ,min(P−1, N K,g)−1do 6:PREFETCHK(g, q, qmodP, s (1),J)▷ ℓ= 1: predicated A loading 7:end for Group-local pipelined mainloop 8:fork c = 0,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a16f832-abcd-4875-bdca-2421050f47ab · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bca8173-c57f-4dec-a62b-8923254aea64 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b08da84-25c3-49ef-860a-93f29aa8e4cd · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebc3e5f2-7465-4512-9f8b-73b9875790cd · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49fd00fc-d098-4dab-b424-1691a1e0d443 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning A VO: Agentic Variation Operators for Autonomous Evolutionary Search
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eab88f6-d7a0-44b4-af9a-40fc86ba8dc1 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f364ec6e-911e-4e10-99a4-791808beaefc · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cdcb772-d14a-47c6-aacd-83df05c840f4 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning ELANA: A Simple Energy and Latency Analyzer for LLMs
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24ccee6c-0cc0-4d8e-94a7-3c3a92efa665 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30b4fb0c-0755-4a82-bcac-3fcdb7c60651 · outbound
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning A Survey on Deep Neural Network Pruning-Taxonomy, Comparison, Analysis, and Recommendations
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.