Pith. sign in

Paper Citation Record · LEDGER

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy

As of 5 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2606.09080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.09080 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T17:02:30.894934Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T08:00:57.256732Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact20
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch12

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8783605d-19ea-4427-9ffc-c4aaf8f94493 · outbound

This paper cites A Deeper Look at Depth Pruning of.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy A Deeper Look at Depth Pruning of

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:4dc8ab3fedf80b8350daff52496ad3330c76a928f89ebca530305d4298c2d330

Observation de5b5d27-1d7b-4d51-b29e-69eec0c6cae3 · outbound

This paper cites V isi P runer: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLM s.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy V isi P runer: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLM s

Reference 2

Resolution
verified exact
doi, observed 2026-06-27T17:11:05.744198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:e9a73914a6ccfe6ee036dabbe27dde296214ca39a6fde99c50e0582c4966c862

Observation b9fbb85d-c413-4312-bfff-6abdd700f1a6 · outbound

This paper cites 2026 , eprint=.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy 2026 , eprint=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:b026bd6dc05f284fb77ac16a91163bed5af44c978d95205cfcd6aa48c68b33e4

Observation 2f11002f-4650-45fb-9a05-9e33b9ff5069 · outbound

This paper cites 2026 , eprint=.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy 2026 , eprint=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:b5c53df39a7141522de6da9b0fc8d6c1222edeca5767a39b3f23904de973b353

Observation 8b879f8a-f1d6-4ea7-b22f-11659bb255f0 · outbound

This paper cites 2026 , eprint=.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy 2026 , eprint=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:66fe1e8b952d533e42713ebac4e6f832ed8d56fdb197459003fb2423ebdd4fb3

Observation c9cb7c60-71f2-432c-9569-ccf1b83cd1a5 · outbound

This paper cites 2025 , eprint=.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy 2025 , eprint=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:b423ba90ae9ccab8246a17094fd431946c0b6bcf5383248f4b27aabac9bde363

Observation d2f842d4-a061-4e05-b439-391fa8fad85f · outbound

This paper cites 2024 , eprint=.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy 2024 , eprint=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:1201b0c943cb41ee6ccfb74dde9ae539f8f82dc22942bb8125140b252154ee5b

Observation 9a1947a6-af24-42f5-84d2-8904b204747c · outbound

This paper cites A Survey on Deep Neural Network Pruning: Taxonomy, Comparison, Analysis, and Recommendations , year=.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy A Survey on Deep Neural Network Pruning: Taxonomy, Comparison, Analysis, and Recommendations , year=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:b493395427178dd948ad996a657999029cf7f566b46cc7a82782cfa578cdea74

Observation d1ef2fcd-5187-423c-88cc-b87f68175912 · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy A Survey on Efficient Inference for Large Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-27T17:11:05.747375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:ae5ec4a333ed198e6302b036b4af9389d14928adf3ab039f61dc6640720107b7

Observation ba36a10c-bb6c-4724-80a4-ed517f74857e · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-27T17:11:05.742050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:77f43fe6f9487ee15886cbe4f86aae214121a17e520bb0706f2f6c256f711625

Observation c0cabbab-a57f-415c-b933-72ef86a254fe · outbound

This paper cites Forty-Second.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Forty-Second

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:1dac49896415a55013cf9ca22e645da7790ced49b4c5e83ba5f08ee9b8166331

Observation 17badc65-825f-4f6e-ae3e-3b20b32e5490 · outbound

This paper cites ShortGPT: Layers in large language models are more redundant than you expect.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy ShortGPT: Layers in large language models are more redundant than you expect

Reference 12

Resolution
verified exact
doi, observed 2026-06-27T17:11:05.731389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:b83586c9f3b7f95bfd29247a65a34f1d612bca47ecbb3f231168be0ca8c8a623

Observation 275aa7b1-8a36-4231-bc2c-6b6ccf9bca99 · outbound

This paper cites Findings of the.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Findings of the

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:5d481699781a5613722da16f1eecbcc4113477586c3a790e6b01bee3f8b2728a

Observation 84c4bc9e-109a-40ae-853e-897eeb02ff08 · outbound

This paper cites Less is More: Towards Green Code Large Language Models via Unified Structural Pruning.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Less is More: Towards Green Code Large Language Models via Unified Structural Pruning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:11:05.739497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:af577fe3265b81346c4457e65a658ed470eb3f933688a527d23422eeb55bbe35

Observation b0d29c0c-4b70-4e9c-b55a-6ddfb3d18aa7 · outbound

This paper cites MultiPruner: Balanced Structure Removal in Foundation Models.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy MultiPruner: Balanced Structure Removal in Foundation Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:11:05.729367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:343994c51c7d99f06a15a0fc42966a7b97b8fbb3487405c71e4e22340c21920b

Observation cdefe15b-5862-4af1-942a-d27eade948f5 · outbound

This paper cites Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:11:05.718176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:c130ddcef874035ff2a3a34f1680de01edaff0e75d1750685e4b131d888e0eb9

Observation 2828d1fd-c1dc-4657-a7cd-c2c0feeb28ff · outbound

This paper cites Efficient.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Efficient

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:10b530b35233cf376a41c8ceea8a6cf33ff3d8535b4f62736d2bdcd32daf7aa7

Observation 6b465e38-2965-4947-bac9-9a1054e87cb6 · outbound

This paper cites and do Nascimento, Marcelo Gennari and Hoefler, Torsten and Hensman, James , year = 2023, month = oct, langid =.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy and do Nascimento, Marcelo Gennari and Hoefler, Torsten and Hensman, James , year = 2023, month = oct, langid =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:fd4b880ae4cd99928b135193485db9c7914be619ee749b5a790f1164956ca32c

Observation abc7e229-8f1b-4c6b-8e99-e5ba17ba664b · outbound

This paper cites Thirty-Seventh.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Thirty-Seventh

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:3baacb8ad290eb8a503988247543f4300d8fd6e915b41da3f4b7e5fe9ca4df5d

Observation 13555b09-9dcc-4d53-aa41-9f72109937b8 · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:96e2dc65df60ed2829634f92e6709159dc32c0fa4f4951c5ee44a5acc4e2061d

Observation 4501e253-9469-4716-b7e5-91172e0ccc03 · outbound

This paper cites arXiv preprint arXiv:2503.09657 , year=.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy arXiv preprint arXiv:2503.09657 , year=

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:11:05.711516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:8d44c80f0b7ce0f172bde39030aa62f55b5be61bca24ce17d786801d8f795fe6

Observation f1225104-a479-4adc-ac84-cd75d72aa7aa · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:11:05.686954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:ba76d9aed0af27f5f2f01d240aae46ff219269e1e7318bd64ac9a734568b30b2

Observation aa8fb56b-a7de-4827-b08b-86e083e06b31 · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:e971940dd1b3b10ed5002f7bcdd27cae0fdb1f500fc02a7fbf71fbff485371f0

Observation cb43ce00-2dc7-48f6-a595-172b0d292c69 · outbound

This paper cites Proceedings of the 40th.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Proceedings of the 40th

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:5b8ac022d1565b74e2bad7d04579be612b40c3e2d84be0733e54cace6751eaa8

Observation 7c851539-8b1b-46bb-98fd-1598d1b7ff71 · outbound

This paper cites Zico , year = 2023, month = oct, langid =.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Zico , year = 2023, month = oct, langid =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:13c60248c9a27a669c16b00b4ae6f6cb861dcaddca1a02ca9b09a1dec557579b

Observation b806d1f4-0a01-4e44-82ec-ee543db4bb41 · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:6f0776bdc0d72ddef03a2fdaa9d921b3b8fb859e681b0003b20b73e52ea1c3c3

Observation eed05c60-e511-4364-9cef-9674d7cfd015 · outbound

This paper cites Prompt-Based.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Prompt-Based

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:743f8da804fc177356a80544260d0da5af37751db59b30d484360d7762c77d16

Observation b95c5a78-460d-4c9f-a925-a46f99c932c6 · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:7dc64c29b69f549c1dfc380131af4fb9c90cd753b45ac16c296bbb9b20aa0426

Observation 11a8c357-df10-4682-bf4f-a0fbbaf4da1e · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:2808f1bd82cdac19d475d621f0ffb23756bec5150ef05452696348d37bc7af7b

Observation 41288731-b927-4803-ab16-29cdfb9b73bf · outbound

This paper cites BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T17:11:05.692659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:0a79cae6cc850fee0bc010dcb6e27b703f54e1b62aa08bd0b4cfa2aff0c57579

Observation cda5dc46-fd2f-429b-b1ae-91a3a90a19f6 · outbound

This paper cites Efficient attention mechanisms for large language models: A survey.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Efficient attention mechanisms for large language models: A survey

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:11:05.721112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:f8a392efa4ae165f363071dfd5985ae7927506f07f965f2525a6816fe03bf3cc

Observation ea86ba86-9faf-4c66-97a1-ed85412abfd5 · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:78052ff128d986f7d29a94f5918b252705f4f118d1076f4dc366160617ce353d

Observation acbc2a59-5538-4655-b371-285df55c6ae0 · outbound

This paper cites Forty-Second.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Forty-Second

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:2f186a655661de9acb294cb464e6355c686ddc8edc4a0112bee64c230c8c0586

Observation 6c4209ae-fede-44b8-8282-b62f093eec0c · outbound

This paper cites A Survey of Large Language Models.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy A Survey of Large Language Models

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T17:11:05.698930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:05090ceea1f7d7741358becdae1b461352e5464a20ea18bbb556c8fbf00bfe23

Observation 2875f540-9ce2-4553-b32b-6c5a12781d87 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Kimi K2.5: Visual Agentic Intelligence

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-27T17:11:05.704357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:f713ec2f31dc37d85e2641948857ae29f49f185843f8ccc442a536675ccc857d

Observation 697f0eca-b0e9-4a95-baac-11bb9e77dca2 · outbound

This paper cites A survey of scientific large language models: From data foundations to agent frontiers.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy A survey of scientific large language models: From data foundations to agent frontiers

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:11:05.701705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:aceb45672a013b98b6fdc7d405f9da30b9f26a5fb13f616d6542a21339549a62

Observation 6918d03b-d60f-4c51-9979-55617cbe470e · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy SGLang: Efficient Execution of Structured Language Model Programs

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T17:11:05.710367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:44c2006f6bac6e87a1621dc38b5266d5a23e665a59b90fc4b071d8cd58727fec

Observation 5c7f8f31-68f9-495d-8fec-bc315910deef · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T17:11:05.696063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:b07f247b10a9caf185d4a7ab4137ab0c1291b35b517c5fe41acf30edd5b460f4

Observation bdc9e7dc-610f-48b6-b142-be9a602932f4 · outbound

This paper cites Int vs fp: A comprehensive study of fine-grained low-bit quantization formats.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Int vs fp: A comprehensive study of fine-grained low-bit quantization formats

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T17:11:05.689953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:1468dded410e747ef33966d377a930249811bf74c7b5347f7f673b3bcc9424b6

Observation be78f15b-4499-4105-8cfc-abc452af05bf · outbound

This paper cites What Matters in Transformers? Not All Attention is Needed.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy What Matters in Transformers? Not All Attention is Needed

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:11:05.719166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:d93eec21779abea5aacb587b2d0ffeea8f6d42b104ec38938c22362b467a5f2f

Observation 49bcd43b-a40e-4fbc-900f-aed73f7be525 · outbound

This paper cites T., and Cox, D.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy T., and Cox, D

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:11:05.736767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:872735ae562c56fbed40532487b4960d915ebe189e708da9072af00ccfe2078b

Observation bf44697e-bde7-46d0-ba20-1691a6e53b75 · outbound

This paper cites B lock P runer: Fine-grained Pruning for Large Language Models.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy B lock P runer: Fine-grained Pruning for Large Language Models

Reference 42

Resolution
verified exact
doi, observed 2026-06-27T17:11:05.723516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:c8a44e03ba65da54591dfcadc69062106b8de6af61a7fe0959ef3b39916b5806

Observation a2b39285-10d9-4d16-abe8-9ae8b345e9de · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:0e30747bdb974dcd5e8d9424c9bd1474c482a3244b37c95f28ca3d1b4b3c779b

Observation 8f79c042-9072-46f4-bce0-60af6340ecdb · outbound

This paper cites The Llama 3 Herd of Models.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy The Llama 3 Herd of Models

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T17:11:05.728245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:561a7e01ce0bef081275afe5b8bf0bdfa86d1c4ad6a70d602e032c32b5084c72

Observation 7e61f776-9bb7-47ad-9685-58d369709c50 · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:73ab2cbe3e254d91a9121e2aaab28c406dbd91baa6705c057e8970c32ecfd7d5

Observation cc47ff91-6408-4755-b24c-0a84db218b39 · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:58eb505b83d7f1cd7e745c4d0ab271440d42f571f17cdd48a8ded4cfe0edb6cb

Observation f3822e68-3b4e-463f-ad87-9354df9aa08b · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:7c05d417df00ecb651e251b1c84a327efc4d6aef0da6a72d36e8f4cdcbf77d6e

Observation 2e7c57b3-d875-4441-a5e1-64245bf1fd95 · outbound

This paper cites doi:10.48550/arXiv.2512.09946 , archiveprefix =.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy doi:10.48550/arXiv.2512.09946 , archiveprefix =

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:11:05.730689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:248691c324b845327162530b4d91222a75ea5ec3325cbfbb2c3561375e2676e8

Observation 3c7d1934-4b1b-4da0-8885-4c6815a02e8b · outbound

This paper cites Pruning as a.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Pruning as a

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:75f695f429bd284eb5b4b717c64f22e6fb23f0799042b4f58cb90a34e35aa812

Observation 1e20dc23-b4c5-4281-8382-2b0e2038482a · outbound

This paper cites Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:11:05.721871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:545815a4e9f739afc1219ce16adc52aa0bca12c4fce3aaccb7cbba096bddc012

Observation 041beb61-d001-4d66-9b0e-e0df04ce510b · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:69b0ca3bb91e02643c123bb79c81dcfc61fa0dd4f1bcee1bebb416a46305fcb1

Observation 0494ec53-2d1e-4e92-b7c4-c7068b9710bc · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:56cdf18698b63d3d17abc8d36b2d505c7cedafd6ef668f5e38c24098fd1a4e5b

Observation 30de5802-2670-4e62-b99d-24638f747198 · outbound

This paper cites and Shen, Yelong and Wallis, Phillip and.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy and Shen, Yelong and Wallis, Phillip and

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:81671690b1b382be920ceaed1628c122a89988c67473bd3c65b6fe64b04f69b3

Observation 58f6ec90-4037-46a0-9c1a-045bcac9556c · outbound

This paper cites Fluctuation-Based.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Fluctuation-Based

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:9358da03e114dce072bb59422744fde4dc6bc0dc1542aeed00d110710b50203e

Observation 76abf642-63b2-44fe-9b82-dd79686ed455 · outbound

This paper cites arXiv.org , howpublished =.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy arXiv.org , howpublished =

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:57e3bc52bff45e7a28ed2d74443849522cfc31799afe128250ae008aa512ff9f

Observation 21ade274-e612-49a2-9f64-dbccceaacb97 · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:7882272a1f6987848710809c1e00584c791760e800a080f2a3b9dcc0ce177279

Observation b41af558-4887-45cd-ae4a-e366127ecc13 · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:aaa692067b6b2f6534ca7b5d1d22ded2a86f9ae283dba3d2686b6b3ce6966bfb

Observation da4f2d62-68cf-4dd7-a7ca-e00f77d5f5b9 · outbound

This paper cites Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:11:05.707747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:b2940aad4789c771b0c1541908e590ac4eeb7c6c6fcc30a15e33ca4f5168f8d3

Observation 8099956c-c11a-4163-88c5-888fc68c228d · outbound

This paper cites doi:10.5281/zenodo.12608602 , url =.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy doi:10.5281/zenodo.12608602 , url =

Reference 59

Resolution
verified exact
doi, observed 2026-06-27T17:11:05.723738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:78e7da6f87810b1eb1760af8f55e732871ccea5bae79fe78bc0e088f9b03d2df

Observation d35400c8-c124-4bba-8c7a-0a005f247639 · outbound

This paper cites Pointer Sentinel Mixture Models.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Pointer Sentinel Mixture Models

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:47:30.322509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:caa91755d02f480d526f281417cb80f4a646d729ae8edd561cc82342eae9ebbe

Observation c7af06e7-606c-443d-8c42-29f798c0fbd4 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:47:30.335240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:4b5f5d7b4bef7c45dfeb6aa1908fa9282945eb41b9385d815392dc93838339f7

Observation 009fd7da-e152-4001-b8bd-1aadff0e4700 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:47:30.331155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:3f8402707c7fcf2a5875cb25c472b84f4d1200e9a60ce7a3c2436842980c1ca2

Observation ea425325-d71a-419e-8183-0abcbbe5973c · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:329be1dcb4301f18b7330095a1a444d5070cb3c7c52c4934cb39e147af336f3d

Observation 7686e2fe-c4ab-4d57-9788-8184fc7db84d · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:47:30.338207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:1f84b60fddab5af70afd9a33385df0bf15dd01dc3e3838825484d93925e7446d

Observation 7a5bd771-d3e8-4cc3-b9a7-4321826b980b · outbound

This paper cites Communications of the ACM , volume=.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Communications of the ACM , volume=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:26fa191e09ab2285294e65cc47e1ef3101b13d4e8e05aa087cdb002f0daf289a

Observation 3ff8f937-c4ac-47db-9c91-ab5b59053f0b · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:47:30.327288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:8f1497acc301a67bada0881f022ebffb1e09221ddcf764456b7d38d45a2bc234

Observation cdaf0ffa-2dd0-4d6a-938e-f5425b47ea38 · outbound

This paper cites an unresolved cited work.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:d9043111ecaf11e596f3c85e799b6b0ca7a6a01050e9d6eba1a6a3c088395f82

Observation 2fe44297-84c4-4251-aea0-8bb23138e7be · outbound

This paper cites Utptrack: Towards simple and unified token pruning for visual tracking.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Utptrack: Towards simple and unified token pruning for visual tracking

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T17:11:05.726523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:84a5c00e4dd42dc549b07e89d610cf19e6614f3eeecca9d0788bb1a40c2143c4

Observation 2e16b766-b2e4-44b0-92c6-bad5ebd4c51d · outbound

This paper cites From Data to Model: A Survey of the Compression Lifecycle in MLLMs , url=.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy From Data to Model: A Survey of the Compression Lifecycle in MLLMs , url=

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:11:05.734110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:7055d3a6a346d985ef4b171ded935eea991f6cafbeed933c35fe297635565cc3

Observation 50933117-11b6-4df4-9c3a-4739228a6881 · outbound

This paper cites 2026 , eprint=.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy 2026 , eprint=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-27T17:02:30.894934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:770920aa6ae58b10f2630ef7f851b92b2228ac4220206df249b21b018b03a6d2

Pith citing papers

Observation 0eab88f6-d7a0-44b4-af9a-40fc86ba8dc1 · inbound

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning cites this paper.

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-07-31T08:00:57.256732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:00:57.256732Z digest=sha256:1fb0e0ce5beaa33d3d3fd3e27ec65f1aa57c200672cc4c2e6ad718c33ddfabd8