Pith. sign in

Paper Citation Record · LEDGER

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning

As of 5 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2607.28418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28418 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T08:00:58.889699Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 14a5e14e-eb41-47ba-ba23-52a3b4de683e · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Check- points.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Check- points

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:55.639927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:55.639927Z digest=sha256:547a67e6ca3e3c69da12a1379f3c34f23ddc751529353aae506c05af30c0ab5b

Observation 28e3d8ed-21e3-46ad-a6e6-ef790d2c7e65 · outbound

This paper cites Aayush Gautam, Mukul Gagrani, Junyoung Park, Mingu Lee, Chiris Lott, and Narasimha Reddy.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Aayush Gautam, Mukul Gagrani, Junyoung Park, Mingu Lee, Chiris Lott, and Narasimha Reddy

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:56.693409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:56.693409Z digest=sha256:97b603e6800a85b6c23b1add26a61e9672e679956daf68f98d9fef40ca41e4e4

Observation 381875ad-a1bb-49c5-90f3-3719cf8d460a · outbound

This paper cites The Llama 3 Herd of Models.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:56.882801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:56.882801Z digest=sha256:5c1545746f401d60d0ad72c6b7aa1b8436699a17559fae324c30fff8391160e7

Observation 10e4d762-f8a0-43a9-91d1-7b22c387d452 · outbound

This paper cites Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:56.974233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:56.974233Z digest=sha256:8a987bb25a484449f8afef15a4985c27e65495325a58fcddfc4dbb858d567636

Observation 98b69cac-ce80-43e4-a590-29f0286678a3 · outbound

This paper cites What Matters in Transformers? Not All Attention is Needed.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning What Matters in Transformers? Not All Attention is Needed

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:57.075282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:57.075282Z digest=sha256:ec0c0d78e44adc19d3ec1d0cad3b876de6bead6b41e176f9e09a500ce7477e19

Observation 1549c1fe-7474-49c3-982e-86d91dba9b22 · outbound

This paper cites SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:57.151768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:57.151768Z digest=sha256:09a7236e50d75dba49e9335e6b3e4788199f61c56e10187d4d64ffea7bf73fec

Observation c213ac60-5bda-4ec8-91fd-8196332a3d9b · outbound

This paper cites Deterministic Differentiable Structured Pruning for Large Language Models.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Deterministic Differentiable Structured Pruning for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:57.357203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:57.357203Z digest=sha256:a3dc7aad9005f57877b40817a58fe06724b4de35b345f70822e0f42a74b5d957

Observation 495e1156-fba4-46a6-b714-4b891f85dcdc · outbound

This paper cites Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:57.433405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:57.433405Z digest=sha256:adac221bdbf215f938b8161f444f1603057cdb30c6e6b3563764fe211d014cee

Observation d7989027-7692-42dd-8eed-787777d7179d · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:57.514150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:57.514150Z digest=sha256:c2679e614f667569d006759593b6e18924645c64ab3bbd25682b54cfc5453df9

Observation 4d3b58ca-6464-4a25-be78-e48ffc7d3e96 · outbound

This paper cites ISBN 979-8-89176-256-5.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning ISBN 979-8-89176-256-5

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:57.669892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:57.669892Z digest=sha256:b3bb8752b1870258d841ddbea23d6959e54ea65cc48d68882e8e2b5ebb01045a

Observation 5606eacd-749e-4cea-95ae-268e8308412e · outbound

This paper cites Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:57.743179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:57.743179Z digest=sha256:1004830231f9384d1fa7140c1b1d0e783290bff45e6db6da67355217b26111ba

Observation 8917b35a-ef5c-47fd-b702-a68d36f06791 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:57.954548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:57.954548Z digest=sha256:31f1e3c46f757d56ae8cbf61c4ef6649f49f9a507048a6a427f53ee89d18bb85

Observation 38989719-78ca-4c0e-b31c-7faf68d8168f · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Kimi K2.5: Visual Agentic Intelligence

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:58.029846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:58.029846Z digest=sha256:f46bdd322bb565d437a2b26aae54cd8f49996661fc3cc762b94677b17440972e

Observation d115edfb-a0ad-44d3-b67d-fd5229029cd3 · outbound

This paper cites an unresolved cited work.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:58.163890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:58.163890Z digest=sha256:47c6d78af82faefb13f807cd249f2a539defd04787168aa3de2bdd2a10a07054

Observation a565838b-d42b-4e62-9de7-fa3f45a0f6da · outbound

This paper cites ISBN 978-1-4503-6719-6.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning ISBN 978-1-4503-6719-6

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:58.302250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:58.302250Z digest=sha256:810e3fa12569627ad0fa2b9f9d665a55f702e148d17c809d7c865256f6cf4e93

Observation 73e7ee6a-a361-482d-84aa-7685fb4d7dcb · outbound

This paper cites From data to model: A survey of the compression lifecycle in mllms.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning From data to model: A survey of the compression lifecycle in mllms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:58.412700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:58.412700Z digest=sha256:4a6bfe4a7f124cc9626ed84589daef567595e97aae837ffc467bd9f625308757

Observation 9f8c2fc2-b59f-4750-9b50-d251b67ba0b6 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:58.486093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:58.486093Z digest=sha256:107b64c7c2416babcbf48d456da295b879b511375f29e293ced97daeaa9e88cb

Observation dc606566-4ed4-4758-9a86-1855bb5c6cdb · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning SGLang: Efficient Execution of Structured Language Model Programs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:58.573833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:58.573833Z digest=sha256:ef0b0a35d10a359ef60cafc03e644b7a606f08c81fc19ffe48b7a0070d67053e

Observation ff557a35-e2cf-4d95-b30c-1011573e922a · outbound

This paper cites BlockPruner: Fine- grained Pruning for Large Language Models.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning BlockPruner: Fine- grained Pruning for Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:58.648132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:58.648132Z digest=sha256:4b61374bebe1cd976ccae86ad14fef61b290198eda28f86750babb3a9c9961d9

Observation fb712a9d-6a89-4271-91eb-9e7755f88ad2 · outbound

This paper cites mk,mnk->mn.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning mk,mnk->mn

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:58.715077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:58.715077Z digest=sha256:7ad2b9e898aa4048bab1b2deef14df3cbe0366bb12c75775260d2e9c856e0e7c

Observation ebed3562-4213-4a99-9420-a259a3cc4ef5 · outbound

This paper cites ,min(P−1, N K,g)−1do 6:PREFETCHK(g, q, qmodP, s (1),J)▷ ℓ= 1: predicated A loading 7:end for Group-local pipelined mainloop 8:fork c = 0,.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning ,min(P−1, N K,g)−1do 6:PREFETCHK(g, q, qmodP, s (1),J)▷ ℓ= 1: predicated A loading 7:end for Group-local pipelined mainloop 8:fork c = 0,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:58.805392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:58.805392Z digest=sha256:c1d2efbc5300049eb7c5763e4a115d5e5a6bb239196f1dd35ddedbc8e1e8532c

Observation 0a16f832-abcd-4875-bdca-2421050f47ab · outbound

This paper cites an unresolved cited work.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Unresolved cited work

Reference 31

Resolution
malformed identifier
no resolver link, observed 2026-07-31T08:00:58.889699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:58.889699Z digest=sha256:097c2f90d32d3fe9b22233a309b66ee354a9b123fd424ba3d809d1aaa75e35ba

Observation 7bca8173-c57f-4dec-a62b-8923254aea64 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:57.839457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:57.839457Z digest=sha256:c935bb6af77f26a6d037d91a4abf078ef1ae94bb4a6f8200d6985da9bb03ada8

Observation 7b08da84-25c3-49ef-860a-93f29aa8e4cd · outbound

This paper cites CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:56.474852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:56.474852Z digest=sha256:e1e59b1b7bbc5ed3b59b4cbdfd2fece36cef6c71c1573054bcedea80e33ef075

Observation ebc3e5f2-7465-4512-9f8b-73b9875790cd · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:56.297010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:56.297010Z digest=sha256:ae9608035cd2c84e10ab3c15dde143c82c8671ad1c290e01705f2b15e031ff9b

Observation 49fd00fc-d098-4dab-b424-1691a1e0d443 · outbound

This paper cites A VO: Agentic Variation Operators for Autonomous Evolutionary Search.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning A VO: Agentic Variation Operators for Autonomous Evolutionary Search

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:55.772555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:55.772555Z digest=sha256:48e28ae9c2e04b61ca271d1dcb0f5a25fe1d18583c6ba64d492827aaea5678a9

Observation 0eab88f6-d7a0-44b4-af9a-40fc86ba8dc1 · outbound

This paper cites Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:57.256732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:57.256732Z digest=sha256:1fb0e0ce5beaa33d3d3fd3e27ec65f1aa57c200672cc4c2e6ad718c33ddfabd8

Observation f364ec6e-911e-4e10-99a4-791808beaefc · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:57.598751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:57.598751Z digest=sha256:6c2a2ac1883f431fc2ea4a8552889c6bc508e3146173599a75f9564cee7bcf04

Observation 1cdcb772-d14a-47c6-aacd-83df05c840f4 · outbound

This paper cites ELANA: A Simple Energy and Latency Analyzer for LLMs.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning ELANA: A Simple Energy and Latency Analyzer for LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:56.072867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:56.072867Z digest=sha256:944ee6e92aa474eeae74ce21d1a3bd0c07a41259ce2db4f0f914b0dd5df1ea89

Observation 24ccee6c-0cc0-4d8e-94a7-3c3a92efa665 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:56.200139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:56.200139Z digest=sha256:445b7a04e4d9c6a9afd79fcb20048e65cce392cf9a359e57eaadfe70881658c5

Observation 30b4fb0c-0755-4a82-bcac-3fcdb7c60651 · outbound

This paper cites A Survey on Deep Neural Network Pruning-Taxonomy, Comparison, Analysis, and Recommendations.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning A Survey on Deep Neural Network Pruning-Taxonomy, Comparison, Analysis, and Recommendations

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:55.931442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:55.931442Z digest=sha256:a7bef2b674377e2d7861054ffacce5c03cc992a9bb2447f25ffd3ba521a4935a

Pith citing papers

No inbound Pith citation observations are available.