Pith. sign in

Paper Citation Record · LEDGER

Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2309.10285.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.10285 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:22:18.874453Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 147484d9-405c-4013-a6b5-272639a833cc · inbound

Pushing the Limits of BFP on Narrow Precision LLM Inference cites this paper.

Pushing the Limits of BFP on Narrow Precision LLM Inference Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T17:22:18.874453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:22:18.874453Z digest=sha256:96759d1477ade74f1531ffe8cb337d737d919360c0e5ab06820b0d12580dd76a

Observation 72d3a5b1-ca63-4a16-bac2-3fb80da30ac2 · inbound

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters cites this paper.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.672927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.672927Z digest=sha256:f5643aa7df529e75de950dc068611c6949b7fc47755ca9ed3c885e761fe61caa

Observation 0664cf4f-2f0f-4f70-805c-12afad15f361 · inbound

RAP: Runtime Adaptive Pruning for LLM Inference cites this paper.

RAP: Runtime Adaptive Pruning for LLM Inference Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:21:35.642487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T13:20:41.739571Z digest=sha256:86f68f908c605d600da372316e40afa1868abf6f0b11bf85c6199f80446da25d

Observation a4826869-d6d3-49e4-adbe-945c7e9bd8b8 · inbound

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs cites this paper.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.552122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:d67675cc8d5b51fc4ebe809877eac0b451a97d4f50488e399a9e8b080af6a479

Observation f690e976-331a-45ab-a40b-0650d905e719 · inbound

Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models cites this paper.

Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T11:47:19.007455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:47:19.007455Z digest=sha256:1155189b13171391094f7f5a009535d2fa28f2f37f8ccc332e51a42b21a2213b

Observation caea12ae-603b-4c34-a5ee-11c1223ddaf5 · inbound

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities cites this paper.

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:10.309428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T09:45:57.201837Z digest=sha256:3aed53c1fc62caebe302d4a28a7a116a813562e95e9d117f0daa12594e9880c1

Observation c139bd58-344c-4240-af5e-218b7f1200cc · inbound

ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs cites this paper.

ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T19:37:43.766024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T19:35:51.992891Z digest=sha256:37a93435d8515481153fc7d8c6ef35462f75e58ce57c9a6b18bc4f493dd40221

Observation cdefe15b-5862-4af1-942a-d27eade948f5 · inbound

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy cites this paper.

Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:11:05.718176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T17:02:30.894934Z digest=sha256:1c4f4fcb949e7dce87712b47fbb1e268bbc25773d32e2245985d92e7180a428e

Observation 3b184a85-9f3c-4509-a578-95e4ef8985b9 · inbound

Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference cites this paper.

Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T00:09:28.755271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:09:28.755271Z digest=sha256:f884f48b4c7e5513a3a4667d6f7e17a07759bb61bb48327584c028ddd967e17e